ls ./ebooks
Two free things. Pick one.
A 9-guide email course on production AI, or the newsletter with three e-books for frontend developers. Both free, both by email, and you can take both.
cat production-ai.md
Production AI in 9 guides
For engineers moving into AI and ML. A free email course — two guides in the first week, then one a week, each a PDF to download and keep.
Fine-tune, RAG, or Just Prompt Better?
The four-bucket test to run before any AI decision.
Most teams fine-tune to fix a problem that was retrieval. Sort thirty failures into four buckets first, and the answer picks itself.
- The four-bucket diagnosis to run before any AI decision
- What each option really costs, in build time and ongoing
- Why fine-tuning on your documents does not teach them
Python for JavaScript Engineers
uv as the npm of Python, plus the traps JS instincts walk you into.
Not a tutorial — a translation layer, plus a warning about the six places your JavaScript instincts will quietly mislead you.
- uv as the npm of Python, and never touching system pip
- Mutable defaults, falsy collections, late-binding closures
- Types and pydantic, which is zod by another name
Prompts as Code
Version, review and test prompts like the load-bearing code they are.
The most load-bearing string in your system is a template literal nobody owns. Version it, review it, and test it like code.
- Prompts as versioned files, with eval diffs in the pull request
- Delimiting untrusted input, and stripping your own delimiters
- Surviving a model deprecation without a rewrite
Evals Before Users
A golden set and regression suite, so your fix doesn’t cause the next bug.
A prompt change that fixes one case quietly breaks four others. Evals are the regression suite that catches it before your users do.
- Building a golden set small enough to maintain
- LLM-as-judge, and how to keep the judge honest
- Running evals in CI without a four-hour pipeline
RAG That Survives Real Documents
Chunking, lineage and retrieval that works on messy real-world docs.
Retrieval demos work on clean markdown. Production hands you scanned PDFs, tables, duplicates and ten years of drift. This is what changes.
- Chunking strategies, and why the fixed-size default fails
- Hybrid search and reranking: when each one earns its cost
- Diagnosing whether a bad answer was retrieval or generation
Vector Databases in Production
Why pgvector is usually enough, and when it isn’t.
Under a million vectors with Postgres already running, pgvector is the answer rather than the compromise. Here is where that stops being true.
- HNSW and the one parameter actually worth tuning
- The post-filter bug that silently returns too few results
- Why every vector needs its embedding model version stored
Agents That Don’t Run Up a Bill
The five limits every agent loop needs.
Cost is quadratic in steps, so doubling the limit roughly quadruples the bill. The runaway loop never errors — it just invoices.
- The five limits every agent loop needs, all of them in code
- Truncating tool results, the biggest single cost lever
- Tiering tools so side effects need a human
From Notebook to Production on GCP
Choosing Cloud Run, Vertex AI or batch, and controlling cost.
Shipping a model is mostly a full-stack problem: an API, a queue, a cache, a UI, and the observability to know when it breaks.
- Vertex AI and Cloud Run: picking the right serving shape
- Cost controls somebody will actually sign off
- What to log, and what it costs you when you don’t
Shipping LLM Features Under Audit
Permissions, logging and what legal will ask.
Nothing in a retrieval pipeline knows who is asking. That is the defect security review finds late, and it is not the only one.
- Filtering retrieval by permissions, before ranking
- Why you cannot fine-tune on data with mixed access
- The audit trail that answers the question actually asked
A wrap-up
The nine guides in one place, and what to build next.