What Actually Breaks When You Put an LLM in Production
It's rarely the model. It's the retry logic nobody wrote, the cost budget nobody set, and the eval suite that didn't exist until after the incident.
AI Architect & Security Professional
Notes on shipping AI Agents, RAG systems, and LLM applications for real-world use, plus some case studies from client work.
It's rarely the model. It's the retry logic nobody wrote, the cost budget nobody set, and the eval suite that didn't exist until after the incident.
Treating the network as an enhancement rather than a requirement changes almost every decision you make, from data storage to how you show errors.
Embed, retrieve, stuff into a prompt: the tutorial makes it sound like a weekend project. The chunking strategy is where weekend projects go to die.
Most AI vendor demos are optimized to look good in a demo. Here's what I actually ask before recommending a client sign anything.
Featured case study
All case studies →~70% reduction
in manual re-keying time across the three highest-volume document types
Review queue, not black box
low-confidence extractions are still reviewed by a person before they hit billing
3 weeks to first production document type
shipped incrementally instead of waiting for full coverage
I take on a small number of engagements alongside the writing: security assessments and threat modeling, plus AI engineering work like RAG systems, agents, and evaluation.