System architecture

The deliverable of an architect isn't code — it's a diagram other people can act on. The hard part is choosing the right altitude: a diagram that shows everything shows nothing. The C4 model fixes this by giving you a zoom control for your system, and the RAG reference architecture below is what one real design looks like at each level.

C4 modelreference architecture RAGgatewayloose coupling

C4: zoom your diagram to the audience

A diagram fails when it mixes altitudes — boxes for "the whole system" next to boxes for "a Python class". The C4 model gives you four fixed zoom levels, and you draw the one your audience needs: Context (the system as one box, and who/what it talks to — for executives), Container (the deployable pieces: apps, services, databases — for engineers), Component (inside one container — for the team building it), and Code (rarely drawn; the IDE shows it). Zoom into a real RAG assistant, one level at a time:

Zoom the same system, level by level — ▶ play it

The discipline is one level per diagram. If you're tempted to add "just one more box" from a deeper level, that's the signal to draw a second diagram instead. A reader should be able to name the audience of any diagram you show them in one glance.

The RAG reference architecture, and why each box earns its place

You just zoomed through a retrieval-augmented generation assistant — the single most common production LLM pattern. At the container level, each box is there to solve one problem: the gateway isolates you from the model provider, the retriever + vector DB ground answers in your own documents, the cache controls cost and latency (from the last page), and a guardrail checks inputs and outputs. Established patterns like this are worth memorising because they're the vocabulary of an architecture review — "put a routing gateway in front" should be a phrase, not a paragraph.

Loose coupling: design the seams, not just the boxes

The most valuable line in an architecture is often the one you don't commit to. Put the model behind a gateway interface and the provider becomes a configuration detail — you can swap Anthropic for a self-hosted model, or route by cost, without touching the rest of the system. That single seam is what made the build-vs-buy decision reversible. Toggle the coupling and watch how far a provider change ripples:

Swap the model provider — how much has to change?

tightly coupled: services call the provider directly · loosely coupled: all calls go through a gateway

⚠️ Traps & honesty: C4 is a convention, not a law — the value is consistent altitude per diagram, whichever notation you use (Mermaid, Excalidraw, draw.io) · the RAG boxes here are a common reference, not the only shape; a simple app may not need a separate gateway or guardrail yet · "loosely coupled" isn't free — a gateway is another hop to run and monitor, justified when provider-swappability or routing is a real requirement, not a hypothetical one · the ripple counts here illustrate the difference, they aren't a metric from your codebase.
Takeaways: draw one C4 level per diagram — Context for execs, Container for engineers, Component for the team — and never mix altitudes · learn the RAG reference architecture as vocabulary: gateway, retriever + vector DB, cache, guardrail, each solving one problem · design the seams: a model behind a gateway interface makes the provider swappable and the big decisions reversible · loose coupling costs a hop, so add it where changeability is a real requirement. Next: data architecture feeds these systems.

Second opinion (taught here — these corroborate): The C4 model · System Design Primer · Anthropic — Building Effective Agents.