AI Infrastructure & Platform Engineering
Self-hosted LLM ops — multi-provider model routing, GPU inference servers, containerized inference engines, real-time cost optimisation. I don't just use AI tools — I build and maintain the entire stack that makes them work, 24/7, on my own hardware.
Infrastructure
Production-grade AI infrastructure — deployed, monitored, and maintained on bare metal hardware.
LiteLLM + OmniRoute proxy layer routing traffic across 6+ providers (OpenRouter, Anthropic, Google, local models). Automatic failover, cost optimisation, and capability-based routing.
llama.cpp, vLLM, DFlash on NVIDIA RTX 3090/4090. VRAM management, batch inference, quantized model deployment, throughput tuning for production workloads.
Full STT → LLM → TTS pipeline. Whisper for speech recognition, CosyVoice for synthesis. Service chaining with latency budgets and real-time processing.
Multi-service Docker Compose stacks with networking, volumes, GPU passthrough, and internal service discovery. 20+ services running on a single host.
Custom log watchdogs, GPU VRAM monitoring, cost tracking, anomaly detection. Automated alerting on service failures and resource exhaustion.
Scripting for infrastructure tasks — config sync, health checks, data processing, API integration. Self-directed learning and continuous improvement.
Portfolio
Real systems deployed and maintained. Not tutorials — production infrastructure.
Autonomous agent with tool access to the entire infrastructure stack. Executes shell commands, manages Docker, browses the web, transcribes audio, controls smart home — fully self-directed operation 24/7.
Stack: Python, MCP, Docker, Telegram, Home Assistant, Proxmox
Complete LLM operations stack: model routing across 6+ providers, GPU inference servers, voice pipeline (STT → LLM → TTS), real-time cost tracking and optimisation. All self-hosted on bare metal.
Stack: LiteLLM, OmniRoute, llama.cpp, vLLM, Whisper, CosyVoice
Full platform of 20+ services: Gitea, Nextcloud, Paperless, Vaultwarden, Penpot, Immich, Paperless-NGX, SearXNG, Meilisearch, RAG API — all containerized, monitored, backed up.
Stack: Docker Compose, Proxmox, TrueNAS, NPM, NFS, Postgres
Personal knowledge graph with 102 MCP tools. Full-text search, semantic recall, code analysis, timeline management, ontology reasoning. The infrastructure for structured memory.
Stack: Bun, MCP, PostgreSQL, semantic search, graph databases
Fail2ban, Vaultwarden password management, HTTPS termination via NPM, automated backups to Proxmox Backup Server, disk cleanup, system hardening, network segmentation.
Stack: Nginx Proxy Manager, Vaultwarden, Proxmox Backup, ZFS
Token economics analysis, model routing by cost/capability, prompt caching strategies, provider failover. Reduced inference costs through intelligent traffic distribution.
Stack: OpenRouter analytics, LiteLLM spend logs, custom routing logic
Get in Touch
Looking for an AI Infrastructure or Platform Engineer? I build and maintain production-grade AI systems — from model routing to GPU inference to full platform operations.
Contact Me