Principles and Systems Architecture for Reasoning, Acting, and Multi-Agent Orchestration
Harvard University · Prof. Vijay Janapa Reddi
Part of the Machine Learning Systems Textbook Series (Volume III)
All active development, open-source code, labs, and draft chapters for Agentic AI Systems are hosted inside Volume III of the unified Machine Learning Systems textbook series.
👉 Read Online: https://mlsysbook.ai
👉 Primary Repository: github.com/harvard-edge/cs249r_book (seebooks/vol3)
👉 Community & Discussion: github.com/harvard-edge/cs249r_book/discussionsPlease star and watch harvard-edge/cs249r_book for the latest chapters and releases.
Traditional machine learning treats models as stateless function approximators ($y = f(x)$). An inference request arrives, compute executes across tensor cores, and tokens or predictions stream back behind glass.
Agentic AI Systems transition machine learning from static inference to dynamic, multi-step execution loops.
In an agentic system, foundation models act as central reasoning units (processors) operating within stateful runtimes. They observe evolving environments, maintain persistent memory hierarchies, formulate plans, invoke external tools, recover from runtime errors, and coordinate with peer agents across networks.
Engineering these systems requires solving classical computer systems problems with a completely new substrate:
- Memory Hierarchies: Working context windows (L1), vector retrieval caches (L2), and durable episodic stores (storage).
- Execution & Scheduling: Interrupts, timeouts, preemption, and speculative trajectory rollouts.
- Fault Tolerance: State checkpointing, self-healing deliberation, and rollbacks.
- Safety & Containment: Sandboxing, capability-based security, and deterministic human-in-the-loop gates.
- Tokenomics & Efficiency: P99 latency tail minimization, KV cache compaction, and cost-aware model routing.
The curriculum follows an 18-chapter progression across systems principles:
| Part | Chapter | Focus |
|---|---|---|
| Foundations | 01. Introduction | From Stateless Inference to Stateful Trajectories |
| 02. The Cognitive Processor | LLMs as CPUs: Context Windows, Instruction Fetch, and Latency | |
| 03. Deliberation | Search, Reflection, Chain-of-Thought, and Rollouts | |
| Memory Architecture | 04. Working Sets | KV Cache Management, Dynamic Slicing, and Context Windows |
| 05. Virtual Memory | RAG, Vector Indexing, Hierarchical Chunking, and Retrieval Paging | |
| 06. Episodic Memory | Durable Multi-Session State, Synthesis, and Long-Term Retention | |
| Runtime Infrastructure | 07. Checkpointing | Trajectory Serialization, Crash Recovery, and Replay Auditing |
| 08. Actuation | Tool Binding, Schema Protocols (MCP), and Environment Bridges | |
| 09. Virtualization | Execution Sandboxes, Containerization, and Blast Radii | |
| 10. Interrupts | Async Callbacks, User Steering, Timeouts, and Signal Traps | |
| 11. Scheduling | Multi-Tenant Agent Queues, Priority Inversion, and Compute Budgets | |
| Learning & Evolution | 12. Data Flywheels | Synthetic Trajectory Generation, Self-Play, and Distillation |
| 13. SFT for Agency | Instruction Tuning on Structured Reasoning and Tool Calling | |
| 14. RLVR | Reinforcement Learning with Verifiable Rewards & Rule Bounds | |
| Scale & Governance | 15. Multi-Agent Systems | Consensus Protocols, Agent-to-Agent IPC, and Delegation Topologies |
| 16. Observability | Distributed Tracing, Telemetry, Eval Harnesses, and Drift Detection | |
| 17. Tokenomics | Unit Economics, Pareto Frontiers, and SLA Optimization | |
| 18. Conclusion | The Future of Autonomous Systems Engineering |
| Resource | Description | Link |
|---|---|---|
| Online Textbook | The full 4-volume textbook series | mlsysbook.ai |
| Agentic Volume | Volume III source files and draft content | cs249r_book/books/vol3 |
| Physical AI Volume | Volume IV (Physical AI & Embodiment) | harvard-edge/physical-ai · cs249r_book/books/vol4 |
| Labs & Simulations | Hands-on labs and systems simulator (MLSys·im) | mlsysbook.ai/labs · mlsysbook.ai/mlsysim |
| Harvard Course | Harvard CS249r Course Portal | harvard-edge.github.io/cs249r_fall2025 |
If you reference this work in academic research, please cite the parent textbook project:
@book{reddi2026mlsys,
title = {Machine Learning Systems: Principles and Practices of Engineering Artificially Intelligent Systems},
author = {Janapa Reddi, Vijay},
year = {2026},
publisher = {MIT Press},
url = {https://mlsysbook.ai}
}