Beneficial prerequisites
These accelerate learning and deepen understanding but aren’t strictly required — the book explains what you need when you get there. Worth prioritizing if you’re architecting multi-agent or enterprise systems rather than single-agent ones. See Essential and Recommended for the tiers above this one.
On this page
Distributed systems concepts
1.3 (multi-agent systems) and 1.5A–1.6 (stateful orchestration) center on multi-agent coordination across processes, servers, or organizations — exactly the setting where assuming instant, perfectly-ordered message delivery causes real bugs. If you’re only building single-agent systems, you can defer this tier entirely.
Synchronous vs. asynchronous messaging, and failure recovery
If three agents are coordinating on a task and the network drops for five seconds, how should the system recover? Can you explain the difference between synchronous and asynchronous messaging?
Video
~7-hour playlist; the same author as Designing Data-Intensive Applications (cited in Recommended > Database Fundamentals). The first 2–3 lectures cover this and the next subtopic.
Eventual consistency, CAP theorem, and consensus
Do you understand what “eventual consistency” means and why distributed systems accept it? At a high level, what problem do consensus protocols like Raft or Paxos solve?
- Continue the same lecture series above, or read the original CAP Theorem explainer for the condensed version
GPU architecture and CUDA basics
Part 7 optimizes LLM inference on GPUs — quantization, batching, and KV-cache tuning all make more sense once you know why GPUs are fast at the operations LLMs need. 4.4 profiles GPU utilization directly. You can learn Part 7’s techniques without this — the book explains what you need — but it moves from recipe to intuition with this background.
Why GPUs are fast for LLM workloads
Can you explain why GPUs excel at matrix multiplication specifically, and what GPU utilization percentages actually indicate?
- NVIDIA CUDA C++ Programming Guide, introduction — read for the architecture rationale, not to become a CUDA programmer
Profiling GPU workloads in practice
Once you understand the “why,” can you read a GPU profiler’s output and identify a memory- vs. compute-bound kernel?
Video
More hands-on than a pure fundamentals video — useful once the “why” above makes sense, less useful before it.
Prompt Engineering Techniques
Part 1 and every framework in Part 2 use prompting; 5.2’s Tree-of-Thought reasoning depends on sophisticated prompt construction. Better prompting skill makes every lab in the book faster. This extends the brief mention in Essential > LLM Fundamentals > Prompting paradigms.
Zero-shot, few-shot, and chain-of-thought prompting
Can you write a few-shot prompt that demonstrates a task to a model, and explain when chain-of-thought prompting is worth the extra tokens?
Video
System messages, templates, and role-based prompting
Do you know the practical difference between a system message and a user message, and can you build a reusable prompt template?
- OpenAI’s prompt engineering guide · free, official, kept current
Debugging a failing prompt
When a prompt doesn’t produce what you expect, do you have a systematic way to isolate why — versus guessing and rephrasing at random?
- Practice: take 10–15 prompts across different tasks, deliberately break each one, and diagnose why before fixing it
Async programming in Python
If you’re only configuring frameworks, async stays hidden. If you’re implementing custom agent logic, it’s essential: 2.9 (streaming) uses async patterns for concurrent tool execution, and Part 7 optimizes throughput with it.
async/await and the event loop
Can you write an async function, use await, and explain what the event loop is actually doing
while your coroutine waits on I/O?
Video
Concurrent execution and when async actually helps
Do you understand the difference between asyncio.gather() and running the same calls
sequentially, and why async helps I/O-bound work but not CPU-bound work?
- Practice: write a script that calls 3–5 APIs concurrently with
asyncio.gather()and time it against calling them one at a time
Threads vs. processes vs. async tasks
Can you explain when you’d reach for a thread, a process, or an async task instead of each other?
CI/CD and DevOps fundamentals
Part 4 teaches CI/CD for agent systems directly, and Part 8 implements monitoring inside those pipelines. Production agents need automated testing, deployment, and rollback — the book provides sufficient guidance if you’re new to this, but you’ll move faster with the basics already in place.
What CI/CD achieves
Can you explain what continuous integration and continuous deployment each solve, and why automated testing has to run before a deploy, not after?
Video
A concrete build/test/deploy pipeline, not just the theory — and containerized, which is how this book deploys everything from Part 4 onward.
Git version control
Can you commit, branch, and merge confidently, including resolving a real conflict rather than avoiding branches to avoid one?
- Git documentation · official, or any interactive git-branching tutorial
Writing a pipeline
Can you write a simple GitHub Actions workflow file from scratch — not copy one and hope?
- GitHub Actions Quickstart · official docs, ~2 hours
Back to Recommended, Essential, or the tier overview.
