I build the
infrastructure that runs AI

Self-hosted LLM ops — multi-provider model routing, GPU inference servers, containerized inference engines, real-time cost optimisation. I don't just use AI tools — I build and maintain the entire stack that makes them work, 24/7, on my own hardware.

The Stack

Production-grade AI infrastructure — deployed, monitored, and maintained on bare metal hardware.

Multi-Provider Model Routing

LiteLLM + OmniRoute proxy layer routing traffic across 6+ providers (OpenRouter, Anthropic, Google, local models). Automatic failover, cost optimisation, and capability-based routing.

LiteLLM OmniRoute OpenRouter

GPU Inference Servers

llama.cpp, vLLM, DFlash on NVIDIA RTX 3090/4090. VRAM management, batch inference, quantized model deployment, throughput tuning for production workloads.

llama.cpp vLLM DFlash

Voice Pipeline

Full STT → LLM → TTS pipeline. Whisper for speech recognition, CosyVoice for synthesis. Service chaining with latency budgets and real-time processing.

Whisper CosyVoice faster_whisper

Container Orchestration

Multi-service Docker Compose stacks with networking, volumes, GPU passthrough, and internal service discovery. 20+ services running on a single host.

Docker Compose Proxmox Nginx Proxy Manager

Observability & Monitoring

Custom log watchdogs, GPU VRAM monitoring, cost tracking, anomaly detection. Automated alerting on service failures and resource exhaustion.

Custom Watchdogs Log Analysis Cost Tracking

Python Automation

Scripting for infrastructure tasks — config sync, health checks, data processing, API integration. Self-directed learning and continuous improvement.

Python YAML Bash

Projects

Real systems deployed and maintained. Not tutorials — production infrastructure.

Hermes AI Agent

Autonomous agent with tool access to the entire infrastructure stack. Executes shell commands, manages Docker, browses the web, transcribes audio, controls smart home — fully self-directed operation 24/7.

Stack: Python, MCP, Docker, Telegram, Home Assistant, Proxmox

AI Infrastructure Stack

Complete LLM operations stack: model routing across 6+ providers, GPU inference servers, voice pipeline (STT → LLM → TTS), real-time cost tracking and optimisation. All self-hosted on bare metal.

Stack: LiteLLM, OmniRoute, llama.cpp, vLLM, Whisper, CosyVoice

Self-Hosted Platform

Full platform of 20+ services: Gitea, Nextcloud, Paperless, Vaultwarden, Penpot, Immich, Paperless-NGX, SearXNG, Meilisearch, RAG API — all containerized, monitored, backed up.

Stack: Docker Compose, Proxmox, TrueNAS, NPM, NFS, Postgres

gbrain — Knowledge Brain

Personal knowledge graph with 102 MCP tools. Full-text search, semantic recall, code analysis, timeline management, ontology reasoning. The infrastructure for structured memory.

Stack: Bun, MCP, PostgreSQL, semantic search, graph databases

Security & Operations

Fail2ban, Vaultwarden password management, HTTPS termination via NPM, automated backups to Proxmox Backup Server, disk cleanup, system hardening, network segmentation.

Stack: Nginx Proxy Manager, Vaultwarden, Proxmox Backup, ZFS

Cost Optimisation Engine

Token economics analysis, model routing by cost/capability, prompt caching strategies, provider failover. Reduced inference costs through intelligent traffic distribution.

Stack: OpenRouter analytics, LiteLLM spend logs, custom routing logic

Let's Work Together

Looking for an AI Infrastructure or Platform Engineer? I build and maintain production-grade AI systems — from model routing to GPU inference to full platform operations.

Contact Me