Blog
Notes on AI engineering, LLMs, RAG and backend systems. The things I build with and think about.
Is an LLM 'intelligent'? How AI differs from a chair or a table
What 'intelligence' and 'understanding' actually mean, where a language model sits between an object and a mind, and why the line matters if you build with these systems.
- AI
- LLMs
- Concepts
How FastAPI actually works: ASGI, async and dependency injection
The ASGI contract that replaced WSGI, how the event loop turns async def into concurrency, and how Starlette, Pydantic and dependency injection fit together over the life of a request.
- FastAPI
- Python
- Backend
Llama 3 70B vs Claude: how to choose an LLM for your product
How to choose between an open-weights model you host yourself and a hosted frontier model. Quality, latency, cost at scale, privacy, and the operational work that rarely makes the pitch deck.
- LLMs
- Architecture
- Cost
Building with Claude and Cursor: an AI engineer's workflow
How AI coding tools fit into my day: where they save time, where they cost it, and the habits around context, prompting and review that decide which one you get.
- Tools
- AI
- Workflow
What RAG really is, and when not to use it
Retrieval-augmented generation end to end: chunking, embeddings, vector search, reranking and grounded prompts. Plus when RAG is the wrong tool.
- RAG
- LLMs
- AI
Model Context Protocol (MCP), explained simply
MCP is an open protocol that lets LLM apps talk to external tools and data through standard servers. The host/client/server model, why it beats one-off integrations, and a worked example.
- MCP
- LLMs
- Tooling
Context caching: how to cut LLM inference cost
How prompt caching reuses computed attention state for repeated prefixes, when it pays off, what the cost maths looks like, which providers support it, and what breaks your cache hits.
- LLMs
- Performance
- Cost
Designing a scalable FastAPI backend
Production patterns for FastAPI under load: staying async end to end, pooling connections, keyset pagination, removing N+1 queries, offloading work to a queue, caching with Redis and scaling out horizontally.
- FastAPI
- Backend
- Architecture
Building LLM agents that don't fall apart in production
The agent loop is simple. Keeping it bounded, deterministic and observable is where the work goes. Guardrails, schema-validated tool calls, evals, and the failure modes that break agents in production.
- Agents
- LLMs
- Production
Vector databases compared: Pinecone vs FAISS vs pgvector
Three ways to store and search embeddings: a managed service, a library and a Postgres extension. How to decide which one your RAG system needs.
- Vector DBs
- RAG
- Databases
On Medium
From Pseudobytes
All posts ↗Writing from the studio I co-founded. Team pieces on shipping AI in production.