Builder, operator2024–PresentLocal-first, zero telemetry
CUDA llama.cpp Open WebUI LocalAI Python
Why Local
Every prompt you send to a cloud LLM is logged, analyzed, and potentially used for training. When you’re asking about architecture decisions or reviewing internal code, that’s not just a privacy concern — it’s a real one.
Modern GPUs are capable enough and models like Llama have made local inference practical. The setup isn’t complicated, and the freedom is worth the tradeoffs.
How It Works
The inference runs on a separate machine with a CUDA GPU — llama.cpp server with quantized GGUF models, mostly Q4_K_M. Good enough for reasoning and code review, not great for anything requiring factual precision. That’s not the use case though.
Open WebUI runs on my main homelab box and points to the llama.cpp server over the network. It’s the interface — clean, familiar, no vendor lock-in.
LocalAI runs as a second option on the main box, capped to one model at a time. It gives me access to a wider range of models without the memory overhead of keeping them all loaded.
CouchDB handles Obsidian sync — my notes stay in sync across devices and feed directly into the Hermes agent context. Nothing goes through a cloud sync service.
How I Actually Use It
Not as a chat toy — as a working tool.
When I’m at a desk, I run OpenCode on my laptop pointing to the local model. It reads code, suggests refactors, and helps me navigate unfamiliar codebases. Everything stays on the network.
When I’m out, my Hermes agent (Sage) runs on the homelab and I talk to it through Matrix. It has tools — terminal access, file reading, web search — so it’s not just a text completion engine. It can actually do things.
What Works
Code review and refactoring suggestions
Drafting technical documents and emails
Explaining unfamiliar codebases
Brainstorming architecture approaches
What Doesn’t
Anything requiring factual accuracy without verification
Long-context tasks beyond the model’s window
Anything you’d trust without reading the output first
The local-first constraint means I can experiment freely. If a model hallucinates or gives bad advice, it doesn’t leak to anyone. That freedom is the whole point.
{"menu":[{"name":"Pages","items":[{"label":"Home","subtitle":"Overview","action":"navigate:/","icon":"page"},{"label":"Career","subtitle":"Timeline & principles","action":"navigate:/career","icon":"page"},{"label":"Projects","subtitle":"All projects","action":"navigate:/projects","icon":"page"},{"label":"About","subtitle":"About Meher","action":"navigate:/about","icon":"page"}]},{"name":"Settings","items":[{"label":"Toggle Theme","subtitle":"","action":"toggleTheme","icon":"theme"}]}],"fuse":{"threshold":0.6,"minMatchCharLength":2,"keys":["label","subtitle","searchableText"]},"projects":[{"title":"Global Payroll Platform","description":"In-house payroll platform across 6 APAC markets, $1.2B annually. I define the technical requirements, coordinate payments integrations, and make sure the engineering teams build what the business actually needs.","tags":["Python","SQL","API Integration","ISO 20022"],"link":"#","image":null,"techStack":["Python","SQL","API Integration","ISO 20022"],"size":"large","domain":"fintech","icon":null,"featured":true},{"title":"Fraud Detection Engine","description":"1.5M payment transactions analyzed. Built automated detection rules that caught the obvious cases — the ones that shouldn't need a human reviewing them. Cut manual validation by 80%. QuickSight dashboards so the non-technical teams could see what was flagged and why.","tags":["Python","QuickSight","Analytics"],"link":"#","image":null,"techStack":["Python","QuickSight","SQL"],"size":"medium","domain":"analytics","icon":null},{"title":"Background Check Revamp","description":"India's background checks were taking four months at the 90th percentile. Candidates would leave. Worked with Legal, Compliance, and Business to renegotiate vendor SLAs, parallelize the workflow, and grow the team from 10 to 45. Cut it to one month, impacted 120K annual hires.","tags":["Operations","Compliance","Process Engineering"],"link":"#","image":null,"techStack":["SQL","QuickSight","VBA"],"size":"medium","domain":"data","icon":null},{"title":"Sovereign Homelab","description":"50 services on bare metal: reverse proxy, DNS, VPN, password manager, photo library, media streaming, AI inference, git CI/CD, smart home, Matrix federation. Full observability via Grafana, Loki, Prometheus, and Alloy. No cloud provider.","tags":["Docker","Linux","Caddy","NetBird"],"link":"#","image":null,"techStack":["Docker","Caddy","NetBird","Vaultwarden"],"size":"large","domain":"infrastructure","icon":null,"featured":true},{"title":"AI Node","description":"Private LLM inference on CUDA GPU. llama.cpp server with quantized models, Open WebUI frontend. Used through OpenCode on my laptop and my Hermes agent on Matrix when I'm out. No cloud APIs, no telemetry.","tags":["CUDA","llama.cpp","Open WebUI"],"link":"#","image":null,"techStack":["Python","CUDA","llama.cpp"],"size":"medium","domain":"ai","icon":null,"featured":true},{"title":"AI & LLM on Your Homelab — A Tutorial","description":"Step-by-step guide to running your own AI inference stack at home. llama.cpp, Open WebUI, Caddy reverse proxy, Matrix integration — everything local, everything yours.","tags":["Tutorial","AI","llama.cpp","Homelab"],"link":"#","image":null,"techStack":["llama.cpp","Open WebUI","Caddy","Docker"],"size":"medium","domain":"ai","icon":null,"featured":false},{"title":"Matrix Citadel","description":"Self-hosted Matrix federation with full voice and video via dual LiveKit servers. Twunnel bridges between two homeservers, JWT auth services, OpenClaw Gateway for AI agent orchestration.","tags":["Matrix","LiveKit","Twunnel"],"link":"https://cloudcitadel.in","image":null,"techStack":["Matrix","LiveKit","Twunnel","OpenClaw"],"size":"medium","domain":"infrastructure","icon":null},{"title":"Media & Photo Stack","description":"Jellyfin streaming with automated library management, subtitles, and analytics. Immich photo library with ML-powered face recognition and reverse image search.","tags":["Jellyfin","Immich","Media Server"],"link":"#","image":null,"techStack":["Jellyfin","Immich","PostgreSQL"],"size":"medium","domain":"infrastructure","icon":null},{"title":"Observability Stack","description":"Full infrastructure telemetry: Grafana dashboards, Loki log aggregation, Prometheus metrics, Alloy collector, plus node, process, and disk health exporters. Monitoring 50 services and bare-metal health in real time.","tags":["Grafana","Prometheus","Loki","Alloy"],"link":"#","image":null,"techStack":["Grafana","Prometheus","Loki","Alloy"],"size":"medium","domain":"infrastructure","icon":null},{"title":"Gitea","description":"Self-hosted git server with CI runner and built-in container registry. This is where the folio lives — code pushed here, the runner builds it, the registry stores the image, and it deploys back to the same machine. Source to production, nothing leaves the box.","tags":["Git","CI/CD","Container Registry"],"link":"https://git.nexusno.de","image":null,"techStack":["Gitea","Docker","CI/CD"],"size":"medium","domain":"infrastructure","icon":null},{"title":"Immich","description":"Photo library with ML-powered face recognition, reverse image search, and automatic duplicate detection. Replaces Google Photos without the telemetry — all the smart features, none of the cloud.","tags":["Immich","ML","Photos"],"link":"#","image":null,"techStack":["Immich","OpenVINO","PostgreSQL"],"size":"medium","domain":"infrastructure","icon":null},{"title":"Home Assistant + Frigate","description":"Smart home automation with AI security camera system. Frigate does hardware-accelerated object detection — people, dogs, cars — on the cameras, not the cloud. Home Assistant ties everything together: lights, sensors, automations, and camera events.","tags":["Home Assistant","Frigate","Smart Home"],"link":"#","image":null,"techStack":["Home Assistant","Frigate","Docker"],"size":"medium","domain":"infrastructure","icon":null},{"title":"SearXNG","description":"Self-hosted metasearch that aggregates results from Bing, Wikipedia, and others without tracking. Replaces the browser search bar — same results, no data broker on the other end.","tags":["SearXNG","Privacy","Search"],"link":"https://search.nexusno.de","image":null,"techStack":["SearXNG","Docker","Valkey"],"size":"small","domain":"infrastructure","icon":null}],"skills":{"tools":["Python","SQL","QuickSight","VBA","Docker","Linux","Kubernetes","Grafana","Prometheus","Loki","Caddy","Matrix"],"standards":["ISO 20022"],"domains":["Payments","Infrastructure","Process Engineering"]},"contact":{"email":"hi@meherchaitanya.com","channels":[{"label":"Email","url":"mailto:hi@meherchaitanya.com","displayText":"hi@meherchaitanya.com","icon":"mail","external":false},{"label":"LinkedIn","url":"https://linkedin.com/in/meherchaitanya","displayText":"meherchaitanya","icon":"linkedin","external":true},{"label":"Matrix","url":"https://matrix.to/#/@meher:hanumara.online","displayText":"@meher:hanumara.online","icon":"matrix","external":true}]}}