AI Engineer · Project path

AI Engineer: build a real AI agent, module by module

Each module combines clear theory, commented code examples and hands-on exercises. Once you finish them, all the deliverables come together into a production-ready agent system.

modules
10
deliverables
10
final project
1
estimated duration
~10 wks

Final project: Nexus Support Agent

A multi-agent customer support system. End-to-end: RAG, episodic memory, tool calling, human-in-the-loop, model routing, observability and continuous evaluation in CI.

Architecture

InputCore AgentsInfrastructure
Chat API (FastAPI)Orchestratorpgvector + RAG
Classifier AgentSupportAgent (ReAct)Redis (session)
MemoryManagerCriticAgentLangSmith (tracing)
GuardrailsEscalationRouterEval pipeline (CI)

Which component each module contributes

  • M1 → LLMClient
  • M2 → PromptLoader
  • M3 → SupportAgent
  • M4 → Orchestrator
  • M5 → RAG+Memory
  • M6 → EscalationRouter
  • M7 → ToolRegistry
  • M8 → ModelRouter
  • M9 → DebugToolkit
  • M10 → Observability

Suggested schedule

One week per module, with progressive integration into the final project.

WeeksPhaseWhat you build
Weeks 1-2FoundationsLLMClient + PromptLoader
Weeks 3-5ArchitectureAgent + Orchestrator + RAG
Weeks 6-7ProductionHITL + Tool Layer
Weeks 8-9AdvancedOptimization + Debug
Week 10LLMOpsObservability + Integration
Share WhatsAppLinkedInX

Module 1 · Phase 1 · LLMs & Prompt Engineering

LLM Fundamentals

Architecture, inference and model selection in production

What is an LLM and how does it generate text?

A Large Language Model is a neural network trained to predict the next token given a context. It doesn't "understand" the way we do: it learned statistical patterns from trillions of tokens of text. Every time it generates a word, it computes a probability distribution over the vocabulary and samples from it.

Analogy

Imagine you completed millions of "finish this sentence" exercises. With enough practice, you develop an intuition for which words tend to follow which. An LLM does something similar, but at massive scale and with far more complex patterns.

The Transformer architecture: what you need to know

You don't need to implement a Transformer, but you do need to understand its practical implications:

Key concepts and their impact in production

  • Self-attention: each token "attends" to every other token in the context. Implication: the model can relate information that is far apart in the text.
  • Context window: the maximum number of active tokens. GPT-4: 128K, Claude 3.5: 200K, Gemini 1.5 Pro: 1M. More context = more cost and latency.
  • Decoder-only: modern models (GPT, Claude, Gemini) only generate; they don't have a separate encoder. They process the whole context every time they generate a token.
  • Tokenization (BPE): text is split into sub-words. "tokenization" can be 3-4 tokens. The real cost of a call depends on the number of tokens, not words or characters.

Inference parameters: the control dial

When you call the API, these parameters determine how the model samples its response:

temperature

0 = deterministic (always the most likely token). 1 = more varied. For agents in production: use 0–0.3. For creative generation: 0.7–1.0.

top_p (nucleus sampling)

Only considers the tokens whose cumulative probability reaches p. top_p=0.9 ignores the least likely 10% of tokens. An alternative to temperature; they aren't used together.

max_tokens

The token limit for the response. It directly affects cost and latency. For structured responses (JSON), a low limit prevents unexpectedly truncated responses.

stop_sequences

The model stops generating when it hits this string. Useful for delimiting outputs: ["</response>", "###"]. More reliable than max_tokens for structured outputs.

Common mistake

Using temperature=0 doesn't guarantee identical outputs. LLMs can vary even at temperature=0 because of differences in hardware and parallelism. For exact reproducibility, store the full input and the output.

Model selection: the most important trade-off

Selection guide for production, 2025

  • Claude 3 Haiku / GPT-4o-mini: classification, routing, simple extraction. ~$0.25/M tokens. Latency: <1s. Use it when errors have low impact.
  • Claude 3.5 Sonnet / GPT-4o: reasoning, generation, complex tool calling. ~$3/M tokens. The ideal balance for most agents in production.
  • Claude 3 Opus / GPT-4-turbo: deep analysis, high-impact decisions. ~$15/M tokens. Only when quality is critical and cost is secondary.
  • Mistral / LLaMA 3 (self-hosted): sensitive data, regulatory compliance, cost at extreme scale. Requires your own infrastructure.
Design principle

70-80% of the queries in a support system are simple. Automatically classifying complexity and using Haiku for simple cases can cut total cost by 60% with no noticeable impact on quality.

LLMClient: the system's base wrapper

The right pattern is not to call the SDK directly from each agent. You create a centralized wrapper that handles automatic retries, structured logging, token counting and model selection.

from anthropic import Anthropic
from tenacity import retry, stop_after_attempt, wait_exponential
from enum import Enum
import structlog, time

log = structlog.get_logger()

class ModelTier(Enum):
    FAST     = "claude-3-haiku-20240307"      # barato y rápido
    STANDARD = "claude-3-5-sonnet-20241022"  # balance ideal
    POWERFUL = "claude-3-opus-20240229"      # máxima calidad

class LLMResponse:
    text: str
    input_tokens: int
    output_tokens: int
    cost_usd: float
    latency_ms: float

class LLMClient:
    def __init__(self):
        self.client = Anthropic()
        self.cost_per_token = {
            ModelTier.FAST:     (0.00025, 0.00125),   # (input, output) por 1K tokens
            ModelTier.STANDARD: (0.003,   0.015),
            ModelTier.POWERFUL: (0.015,   0.075),
        }

    @retry(stop=stop_after_attempt(3),
           wait=wait_exponential(multiplier=1, min=1, max=10))
    def call(
        self,
        messages: list[dict],
        model: ModelTier = ModelTier.STANDARD,
        temperature: float = 0.3,
        max_tokens: int = 1024,
        trace_id: str = None,
    ) -> LLMResponse:
        start = time.time()

        response = self.client.messages.create(
            model=model.value,
            messages=messages,
            temperature=temperature,
            max_tokens=max_tokens,
        )

        latency = (time.time() - start) * 1000
        cost = self._calculate_cost(model, response.usage)

        # Log estructurado para observabilidad
        log.info("llm_call",
            trace_id=trace_id,
            model=model.value,
            input_tokens=response.usage.input_tokens,
            output_tokens=response.usage.output_tokens,
            cost_usd=round(cost, 6),
            latency_ms=round(latency, 1),
        )

        return LLMResponse(
            text=response.content[0].text,
            input_tokens=response.usage.input_tokens,
            output_tokens=response.usage.output_tokens,
            cost_usd=cost,
            latency_ms=latency,
        )

    def _calculate_cost(self, model, usage) -> float:
        inp, out = self.cost_per_token[model]
        return (usage.input_tokens/1000*inp) + (usage.output_tokens/1000*out)

The @retry decorator from tenacity automatically handles rate limits (429) and transient API errors with exponential backoff. Without it, any transient failure breaks the agent's flow.

Resources

anthropic-sdk docs, tenacity, structlog, tiktoken

Exercise 1

Temperature benchmark

Understand empirically how temperature affects the output before choosing the value for production.

  1. Write a prompt that asks to classify the sentiment of a sentence (positive/negative/neutral)
  2. Run the same call 10 times with temperature=0. Does it always give the same result?
  3. Repeat with temperature=0.5 and temperature=1.0. What changes?
  4. Record the cost and latency of each call. Does temperature affect cost?
  5. Conclusion: which temperature would you choose for the project's intent classifier?
Exercise 2

Token counting and cost estimation

Before designing the system, know how much each call will cost.

  1. Write the support agent's system prompt (initial draft, ~200 words)
  2. Use tiktoken to count how many tokens it takes
  3. Simulate 1000 five-turn conversations: calculate the total cost with Haiku vs Sonnet
  4. What percentage of the cost comes from the system prompt vs the history?
  5. Document which model you would choose for the intent classifier, and why

Module deliverable

Goes into the final project: LLMClient wrapper. A production-ready Python class that wraps every API call. Every agent in the project will use this wrapper, never the SDK directly. The LLMClient is the base layer of the system. In M8 (Model Router) it will be extended to select the model dynamically per query, instead of receiving it as a fixed parameter.

Module 2 · Phase 1 · LLMs & Prompt Engineering

Advanced Prompt Engineering

System prompts, constraints, controlled CoT and versioning

The system prompt is the agent's constitution

The system prompt is not an "initial instruction". It is the document that fully defines who the agent is, what it can do, how it should behave and when it should ask for help. An agent without a well-structured system prompt is an unpredictable agent.

6-section structure (all required)

  • IDENTITY: name, purpose, personality. Defines the "who" of the agent.
  • CAPABILITIES: an explicit list of what it can and CANNOT do. The "cannots" matter as much as the "cans".
  • CONTEXT: dynamic variables from the environment: user, state, available tools.
  • BEHAVIOR RULES: how to act in specific situations, with explicit edge cases.
  • OUTPUT FORMAT: the exact structure, length and channel of the response.
  • ESCALATION: exact criteria for handing off to a human.
Golden rule

Whatever isn't explicit in the system prompt, the model makes up. Every expected behavior has to be specified. Ambiguity in the prompt is the source of 80% of agent bugs.

Controlled Chain-of-Thought: separating reasoning from the answer

CoT (Chain-of-Thought) improves reasoning quality, but in production we don't want to show the internal process to the user. The right pattern is to separate the two:

Uncontrolled CoT

  • The user sees the internal reasoning
  • It exposes logic that can be manipulated
  • It needlessly increases output tokens
  • It makes the final answer harder to parse

CoT with a separate <thinking>

  • The reasoning stays in internal logs
  • It allows debugging without exposing it to the user
  • The final output is clean and parseable
  • You can monitor the quality of the reasoning

Few-shot with negative examples

Positive examples teach the expected behavior. Negative examples are just as critical: they show the model exactly what to avoid. Without them, the model can fall into "reasonable but wrong" answers.

Common anti-pattern

Including only positive examples in the few-shot. The model learns "what to do" but not "what NOT to do". The most frequent edge cases and failures should appear as explicit negative examples.

Prompts as code: versioning and testing

A prompt that changes without control is a silent regression. The same discipline we apply to code applies to prompts:

Prompt deployment pipeline

  • A git PR with the prompt change + the rationale in the description
  • Automatic evaluation in CI against the base test set
  • Deploy to staging → 10% of traffic → 48h of monitoring
  • If metrics are OK → promote to 100%. If they degrade → automatic rollback

Complete system prompt with controlled CoT

IDENTITY:
Eres SupportBot, asistente de atención al cliente de Nexus.
Objetivo: resolver consultas de soporte con empatía y precisión.
Tono: cercano, claro, sin jerga técnica innecesaria.

CAPABILITIES:
✓ Puedes: get_order_status, create_ticket, send_notification, schedule_callback
✗ NO puedes: modificar precios, eliminar cuentas, acceder a datos de pago

CONTEXT:
Usuario: {{user_name}} | Plan: {{plan_name}} | Estado: {{account_status}}
Canal: {{channel}} | Herramientas: {{available_tools}}

BEHAVIOR RULES:
- Lenguaje agresivo → desescalar sin confrontar: "Entiendo tu frustración,
  mi objetivo es encontrar una solución que funcione para ti."
- Solicitud fuera de alcance → explicar límite + ofrecer alternativa real
- Input ambiguo → preguntar UNA cosa antes de actuar
- Señal de crisis o urgencia alta → escalar a humano inmediatamente

ESCALATION:
Transferir SIEMPRE cuando: ticket_priority="critical" OR usuario solicita
hablar con persona OR confidence_score < 0.70

REASONING FORMAT:
Antes de responder, razona en <thinking>:
1. ¿Qué pide exactamente el usuario?
2. ¿Qué información tengo vs qué me falta?
3. ¿Qué regla de comportamiento aplica?
4. ¿Debo escalar o puedo resolver?
El contenido de <thinking> NO se muestra al usuario.

OUTPUT FORMAT:
- Máx 3 oraciones por turno (canal: chat/WhatsApp)
- Cuando uses herramienta: responde SOLO JSON válido sin texto adicional:
  {"action": "<tool>", "params": {...}, "reason": "<1 oración>", "confidence": 0.0-1.0}

PromptLoader: centralized version management

import os
from pathlib import Path
from jinja2 import Template

PROMPTS_DIR = Path("prompts")

class PromptLoader:
    def get(self, agent: str, version: str, context: dict) -> str:
        """Carga un prompt versionado e inyecta variables de contexto."""
        path = PROMPTS_DIR / agent / f"v{version}_system.txt"
        template_str = path.read_text(encoding="utf-8")
        return Template(template_str).render(**context)

    def latest(self, agent: str) -> str:
        """Lee la versión actual desde el archivo VERSION del agente."""
        version_file = PROMPTS_DIR / agent / "VERSION"
        return version_file.read_text().strip()

    def load_latest(self, agent: str, context: dict) -> str:
        """Atajo: carga siempre la versión más reciente."""
        return self.get(agent, self.latest(agent), context)

# Uso en un agente:
loader = PromptLoader()
system_prompt = loader.load_latest("support_agent", {
    "user_name": "Ana López",
    "plan_name": "Pro",
    "account_status": "active",
    "channel": "whatsapp",
    "available_tools": "[get_order_status, create_ticket]",
})

Resources

jinja2, jsonlines, Anthropic prompting guide, Learn Prompting

Exercise 1

Guided construction of the system prompt

Write the complete system prompt for SupportBot following the 6-section structure.

  1. Write a first version with no structure, just whatever comes to you naturally
  2. Evaluate it: which section is missing? Is any rule ambiguous?
  3. Rewrite it using the 6 sections. Add at least 3 specific BEHAVIOR RULES
  4. Test the prompt by sending 5 edge-case messages: aggressive input, impossible request, ambiguous input, request for sensitive data, and a normal valid query
  5. Adjust the rules based on the results and document what changed in CHANGELOG.md
Exercise 2

Few-shot with negative cases

The most valuable few-shot includes examples of what NOT to do, not just what to do.

  1. Identify the 3 most common kinds of mistakes SupportBot could make (e.g. assuming before asking, answering out of scope, using the wrong tone)
  2. For each mistake, write a pair (user message → WRONG bot response)
  3. Then write the CORRECT response for the same message
  4. Add these 3 negative pairs to the system prompt and test again with the 5 messages from EX1
  5. Did the behavior improve? In which cases?

Module deliverable

Goes into the final project: Prompt Library v1.0 + PromptLoader. Versioned system prompts for the project's 3 agents (support, classifier, critic), with dynamic variables, a base test set and a loader that injects context at runtime. The PromptLoader is used by every agent. In M10, the eval pipeline will hook into it to run the test set automatically on every PR that modifies a prompt.

Module 3 · Phase 2 · Agents & Memory

Agent Patterns

Planner/Executor, ReAct loop, Tool-using and Critic

What is an LLM agent?

An agent is a system where the LLM doesn't just generate text: it also decides which actions to execute, observes the results of those actions, and decides what to do next. The difference from a simple chatbot is that the agent has agency: it can act on the world.

Analogy

A chatbot is like an employee who can only give verbal answers. An agent is like an employee who can also open systems, send emails, create tickets and look up information, all in response to what the customer needs.

The ReAct pattern: Reason + Act

ReAct is the most widely used pattern in production. On each turn, the agent (1) reasons about the current state, (2) decides on an action, (3) observes the result, and repeats until it has enough information to answer the user.

The ReAct cycle, step by step

  • Thought: "The user is asking about their order. I need to call get_order_status with their ID."
  • Action:get_order_status(order_id="ORD-123")
  • Observation:{"status": "in transit", "eta": "tomorrow 14:00"}
  • Thought: "I have the information I need. I can answer the user."
  • FINISH: "Your order is on its way and will arrive tomorrow before 2:00 p.m."
Infinite loop: the most common risk

Without a hard MAX_ITERATIONS limit, an agent can cycle indefinitely if a tool keeps failing or the reasoning doesn't converge. This limit belongs in the orchestrator, not in the prompt.

Critic loop: the agent evaluates its own output

The Critic is a second agent (or a second LLM call) that evaluates the main agent's response before sending it to the user. It answers PASS/FAIL + a reason. It doubles the cost but significantly raises quality in high-impact cases.

When to use a Critic loop

Use it when the cost of a wrong answer is higher than the cost of the extra call. For support: when the agent is about to create a ticket or send a notification. Don't use it on every response, only on actions with side effects.

from dataclasses import dataclass, field
from enum import Enum

class AgentAction(Enum):
    FINISH   = "FINISH"
    ESCALATE = "ESCALATE"
    TOOL     = "TOOL"

@dataclass
class AgentThought:
    reasoning: str          # contenido del <thinking>
    action: AgentAction
    tool_name: str | None = None
    tool_params: dict      = field(default_factory=dict)
    final_answer: str | None = None
    confidence: float       = 1.0

class SupportAgent:
    max_iterations = 8

    def __init__(self, llm_client, tool_registry, prompt_loader):
        self.llm    = llm_client
        self.tools  = tool_registry
        self.loader = prompt_loader

    def run(self, user_message: str, context: dict) -> AgentResult:
        system = self.loader.load_latest("support_agent", context)
        history = []

        for i in range(self.max_iterations):
            # Paso 1: REASON — el agente piensa qué hacer
            messages = self._build_messages(system, user_message, history)
            response = self.llm.call(messages, trace_id=context["trace_id"])
            thought  = self._parse_thought(response.text)

            # Paso 2: verificar stopping criteria
            if thought.action == AgentAction.FINISH:
                return AgentResult(answer=thought.final_answer, iterations=i+1)

            if thought.action == AgentAction.ESCALATE:
                return AgentResult(escalate=True, reason=thought.reasoning, iterations=i+1)

            # Paso 3: ACT — ejecutar la herramienta
            observation = self.tools.execute(thought.tool_name, thought.tool_params)

            # Paso 4: OBSERVE — agregar al historial
            history.append({"thought": thought, "observation": observation})

        # MAX_ITERATIONS alcanzado → siempre escalar, nunca lanzar excepción
        return AgentResult(escalate=True, reason="max_iterations_reached", iterations=self.max_iterations)


class CriticAgent:
    def evaluate(self, agent_result: AgentResult, original_query: str) -> CriticVerdict:
        """Evalúa si la respuesta del agente es correcta antes de enviarla."""
        prompt = f"""
Evalúa esta respuesta de soporte:
Consulta original: {original_query}
Respuesta del agente: {agent_result.answer}

Responde SOLO con JSON:
{{"status": "PASS" o "FAIL", "reason": "...", "suggestion": "..."}}
"""
        response = self.llm.call([{"role": "user", "content": prompt}],
                                  model=ModelTier.FAST)  # Haiku para el critic = más barato
        return CriticVerdict(**json.loads(response.text))

Resources

ReAct paper (Yao 2022), Anthropic tool use, pydantic v2

Exercise 1

Implement the ReAct loop from scratch

Before using the base code, understand the pattern by implementing it yourself with a simple case.

  1. Create a fake tool get_weather(city) that returns hardcoded JSON
  2. Implement the ReAct loop in ~30 lines: reason → parse action → execute → observe → repeat
  3. Test with "What's the temperature in Madrid?": the agent should call the tool
  4. Now test with "Tell me a joke": the agent should finish in 1 iteration without using a tool
  5. Force the infinite loop: make get_weather always return an error. Does MAX_ITERATIONS work?
Exercise 2

Build the CriticAgent and test how effective it is

Evaluate whether the Critic actually improves the quality of the system.

  1. Implement the CriticAgent with the prompt from the base code
  2. Generate 10 SupportAgent responses to a variety of queries
  3. Evaluate each one with the Critic. How many pass? How many fail, and why?
  4. For the ones that fail: is the Critic right? Are there false positives?
  5. Measure the Critic's additional cost: how much does it add per query? Is it worth it?

Module deliverable

Goes into the final project: SupportAgent Core + CriticAgent. The main agent with a ReAct loop, MAX_ITERATIONS, stopping criteria and escalation. Plus a CriticAgent that evaluates high-impact responses before they are sent. The SupportAgent is the engine of the system. In M4 it will be wrapped by the Orchestrator. The CriticAgent will connect to the M10 evaluation pipeline to measure quality in production.

Module 4 · Phase 2 · Agents & Memory

Multi-Agent Orchestration

Orchestrator, classifier, typed handoffs and global timeout

Hub-and-spoke: the most robust pattern for production

A central orchestrator receives every message, classifies it with a lightweight agent (cheap and fast), and delegates to the right specialized agent with the full context.

Analogy

Like a hospital receptionist: they don't make the diagnosis, but they know exactly which specialist to send you to. The classifier is the receptionist: fast, cheap, and with a routing criterion.

The classifier is the most critical piece of the system

  • Use the cheapest model: Haiku with a 5-line prompt classifies better than Sonnet with an ambiguous prompt
  • Exhaustive categories: every query has to land in some category, so include "GENERAL/OTHER"
  • Typed output: the classifier never returns free text; it returns an enum with the category
  • Safe fallback: if the classifier fails, the system routes to the general agent and never breaks
Anti-pattern: handoff without context

The most common mistake in multi-agent systems: agent B receives the user's message but doesn't know what agent A did. The handoff must include the full history, the action already taken and the reason for the transfer.

class Intent(Enum):
    ORDER_STATUS  = "order_status"
    CREATE_TICKET = "create_ticket"
    ESCALATE      = "escalate"
    GENERAL       = "general"

@dataclass
class AgentHandoff:
    """Contexto completo que pasa entre agentes en un handoff."""
    user_id: str
    user_message: str
    intent: Intent
    conversation_history: list[dict]
    previous_actions: list[str]   # qué ya intentó el agente anterior
    context: dict                  # datos del usuario (plan, status, etc.)
    trace_id: str

class Orchestrator:
    global_timeout = 30  # segundos — nunca un workflow dura más

    def handle(self, user_id: str, message: str) -> OrchestratorResponse:
        context = self.context_builder.build(user_id)
        trace_id = self._new_trace_id()

        # 1. Clasificar intención con modelo barato (Haiku)
        intent = self.classifier.classify(message, context)

        # 2. Construir handoff con contexto completo
        handoff = AgentHandoff(
            user_id=user_id, user_message=message, intent=intent,
            conversation_history=self.session.get_history(user_id),
            previous_actions=[], context=context, trace_id=trace_id
        )

        # 3. Routing al agente correcto con timeout global
        with timeout(self.global_timeout):
            agent = self.router[intent]
            result = agent.run(handoff)

        # 4. Evaluar si escalar antes de responder al usuario
        verdict = self.escalation_router.evaluate(result, context)
        if verdict.should_escalate:
            return self._escalate(handoff, verdict.reason)

        return OrchestratorResponse(message=result.answer, trace_id=trace_id)

Resources

signal (timeout), claude-3-haiku, pydantic

Exercise 1

Design the routing scheme

Before implementing, design the complete map of intents and agents.

  1. List every possible query a support user might have (at least 15)
  2. Group them into categories. How many agents do you really need?
  3. Write the classifier prompt with all the categories
  4. Test the classifier with the 15 queries. Does it classify them correctly?
  5. Adjust until you reach >90% accuracy on the 15 queries
Exercise 2

Simulate a failed handoff

Understand what happens when the handoff context is incomplete.

  1. Implement a minimal handoff: it only passes the user's message, with no history or context
  2. Test with a user who is resuming an earlier conversation
  3. Does agent B "know" what agent A did? Does it answer correctly?
  4. Add the full history to the handoff and repeat. Does it improve?
  5. Document which AgentHandoff fields are essential

Module deliverable

Goes into the final project: Orchestrator + ClassifierAgent. An orchestrator with routing, a global timeout and typed handoffs. A ClassifierAgent on Haiku that routes queries correctly. The Orchestrator is the API's entry point. In M6 the EscalationRouter is added on its output, and in M7 the ToolRegistry is injected into the SupportAgent that the Orchestrator coordinates.

Module 5 · Phase 2 · Agents & Memory

Memory, Context and RAG

Embeddings, vector store, retrieval and memory strategies

The memory problem in LLMs

By default, an LLM remembers nothing between sessions. Every API call is stateless. For a support agent, that's a problem: users shouldn't have to repeat their issue in every interaction.

The 4 types of memory and when to use each

  • Short-term (active window): the current turn's history in the LLM's context. Free in terms of cost, lost when the session closes.
  • Long-term (vector store): the domain knowledge base. Semantic search by embedding similarity. For documentation, FAQs, policies.
  • Episodic (interaction history): what the user said in previous sessions. A structured database with timestamps.
  • Semantic (user entities): persistent data: plan, preferences, ticket history. Structured DB.
The human support agent analogy

Short-term = what they remember from this call. Long-term = the support manual they looked up. Episodic = notes from previous calls with this customer. Semantic = the customer's file with their details and plan.

The RAG pipeline: how it works in production

The 4 steps of the pipeline

  • Ingestion: document → chunking (512 tokens, 10% overlap) → embedding → vector store + metadata
  • Retrieval: query → embed → ANN search (top-10) → metadata filter → reranking → top-3
  • Augmentation: retrieved chunks → inject into the LLM's context
  • Evaluation: does the answer use the chunks? Were the chunks relevant?
Chunking is the most underrated decision

Chunks that are too small (< 200 tokens) lose context. Chunks that are too large (> 1500 tokens) introduce noise. Experiment with your specific domain: there is no universally optimal size.

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.vector_stores.postgres import PGVectorStore
import redis

class MemoryManager:
    def __init__(self, pg_conn_str: str, redis_url: str):
        self.vector_store = PGVectorStore.from_params(pg_conn_str, embed_dim=1536)
        self.session      = redis.from_url(redis_url)
        self.index        = VectorStoreIndex.from_vector_store(self.vector_store)

    def get_context(self, query: str, user_id: str) -> MemoryContext:
        return MemoryContext(
            # Short-term: historial de la sesión actual
            short_term = self._get_session(user_id),

            # Long-term: knowledge base del dominio (RAG)
            long_term  = self._retrieve_relevant(query, k=3),

            # Episodic: últimas 3 interacciones del usuario
            episodic   = self._get_recent_episodes(user_id, n=3),
        )

    def _retrieve_relevant(self, query: str, k: int) -> list[str]:
        retriever = self.index.as_retriever(similarity_top_k=k*3)  # más para re-rankear
        nodes = retriever.retrieve(query)
        # Reranking: ordenar por relevancia real, no solo similitud vectorial
        reranked = sorted(nodes, key=lambda n: n.score, reverse=True)[:k]
        return [n.text for n in reranked]

    def _get_session(self, user_id: str) -> list[dict]:
        raw = self.session.get(f"session:{user_id}")
        return json.loads(raw) if raw else []

    def save_turn(self, user_id: str, user_msg: str, agent_response: str):
        history = self._get_session(user_id)
        history.append({"user": user_msg, "agent": agent_response})
        self.session.setex(f"session:{user_id}", 3600, json.dumps(history))  # TTL 1h

Resources

llama-index, pgvector, redis-py, Cohere Rerank, RAGAS (eval)

Exercise 1

Experiment with chunking

Chunk size is the variable with the biggest impact on RAG quality.

  1. Take 5 support documents (FAQs, guides, policies)
  2. Index them with chunk_size=256 tokens
  3. Ask 10 questions about the content. What % does it answer correctly?
  4. Re-index with chunk_size=512 and chunk_size=1024. Repeat the questions
  5. Which size gives the best results for your domain? Why?
Exercise 2

Measure the impact of episodic memory

Quantify whether episodic memory improves the real experience.

  1. Simulate a two-session conversation: in the first, the user reports a problem. In the second, they come back with the same problem
  2. Test without episodic memory: does the agent remember the earlier context?
  3. Turn on episodic memory and repeat. Does the agent respond differently?
  4. Measure the extra cost of including the episodic history in the context
  5. Is it worth it? Document the decision in an ADR

Module deliverable

Goes into the final project: MemoryManager + RAG Pipeline. A complete ingestion and retrieval pipeline, plus a MemoryManager class that manages the agent's 3 types of memory. The MemoryManager is injected into the Orchestrator. Before each call to the agent, the system retrieves relevant context (RAG + episodic) and adds it to the prompt dynamically.

Module 6 · Phase 3 · Production & Integration

Human-in-the-Loop

Escalation criteria, cascading fallbacks and circuit breaker

Human-in-the-loop is not an edge case: it's design

The most common mistake is treating escalation as something exceptional. In production, 10-30% of interactions will end up with a human. The system has to be designed for this from the start, not have it bolted on later.

Design principle

Define the escalation criteria BEFORE going to production. If you define them once incidents are already happening, you're choosing them under pressure and without data. The criteria should be configurable per environment and measurable on the dashboard.

5 types of escalation criteria

  • Business threshold: ticket_priority="critical", account_type="enterprise"
  • Low confidence: confidence_score < 0.70 on the action to take
  • Out of scope: the agent can't resolve the request
  • Explicit request: the user asks to talk to a person
  • Emotional signal: crisis, extreme urgency, accumulated frustration

Cascading fallback: the system never dies

A production system must always respond, even when everything fails. The cascading fallback pattern defines a chain of gradual degradation:

Fallback chain

  • Level 1: SupportAgent with Sonnet (normal)
  • Level 2: SupportAgent with Haiku (faster and cheaper if there is latency)
  • Level 3: hardcoded generic response + automatic escalation to a human
  • Level 4: friendly error message with an automatically created ticket number
from pybreaker import CircuitBreaker, CircuitBreakerError

# Circuit breaker por herramienta — evita cascada de fallos
ticket_breaker = CircuitBreaker(fail_max=5, reset_timeout=60)

class EscalationRule:
    name: str
    check: callable  # función que recibe (result, context) → bool
    reason: str

class EscalationRouter:
    rules: list[EscalationRule] = [
        EscalationRule("critical_ticket",
            lambda r, ctx: ctx.get("ticket_priority") == "critical", "ticket_critico"),
        EscalationRule("low_confidence",
            lambda r, ctx: r.confidence < 0.70, "confianza_baja"),
        EscalationRule("user_requested",
            lambda r, ctx: ctx.get("user_requested_human", False), "usuario_solicito"),
    ]

    def evaluate(self, result, context) -> Verdict:
        for rule in self.rules:
            if rule.check(result, context):
                audit_log.record("escalation", rule=rule.name)
                return Verdict(should_escalate=True, reason=rule.reason)
        return Verdict(should_escalate=False)

class FallbackChain:
    def run(self, handoff: AgentHandoff) -> AgentResult:
        try:
            return self.support_agent.run(handoff)         # Nivel 1: normal
        except (TimeoutError, RateLimitError):
            try:
                return self.support_agent_fast.run(handoff)  # Nivel 2: modelo barato
            except Exception:
                return self._static_fallback(handoff)        # Nivel 3: respuesta fija

    def _static_fallback(self, handoff) -> AgentResult:
        ticket_id = self._create_fallback_ticket(handoff)
        return AgentResult(
            answer=f"Estamos experimentando problemas técnicos. Creamos el ticket #{ticket_id} y un agente te contactará pronto.",
            escalate=True, reason="system_fallback"
        )

Resources

pybreaker, structlog, pydantic

Exercise 1

Define and test the escalation criteria

Poorly defined criteria produce too many or too few escalations, and both are costly.

  1. Define 5 escalation criteria for SupportBot. Write them as exact conditions
  2. Create 10 test scenarios: 5 that should escalate, 5 that shouldn't
  3. Implement the EscalationRouter and run the 10 scenarios
  4. How many false positives (escalates when it shouldn't)? False negatives?
  5. Tune the thresholds until you get 0 false negatives (the priority) and under 10% false positives

Module deliverable

Goes into the final project: EscalationRouter + FallbackChain. A complete safety module: 5 configurable escalation criteria, a 3-level cascading fallback, a circuit breaker and an audit log. The EscalationRouter connects to the Orchestrator's output. In M10, the escalation rate becomes a business metric on the observability dashboard.

Module 7 · Phase 3 · Production & Integration

Tool Layer & External APIs

Typed wrappers, idempotency, state management and tool registry

The agent never touches infrastructure directly

The most important rule of the tool layer: the agent calls contracts (typed schemas), not implementations. This lets you change the underlying implementation without touching the agent, and test the agent with mocks without real infrastructure.

Anatomy of a well-designed tool

  • Typed schema: a Pydantic model with validation, descriptions and constraints
  • Dry-run mode: validate without executing side effects, so you can check before acting
  • Per-tool timeout: each tool has its own SLA, not the workflow's global one
  • Idempotency: running the same tool twice with the same params = the same result
  • Audit log: every execution is recorded, successful and failed
The LLM can pass invalid parameters

The model can generate out-of-range params, wrong types, or empty required fields. Never trust the LLM's output without validation. Pydantic raises a ValidationError before the action reaches the service.

from pydantic import BaseModel, Field
from typing import Literal

# 1. Schema tipado — lo que el LLM ve y debe rellenar
class CreateTicketParams(BaseModel):
    user_id:  str            = Field(description="ID único del usuario")
    subject:  str            = Field(min_length=5, description="Asunto del ticket")
    priority: Literal["low","medium","high","critical"]
    category: str            = Field(description="Categoría: billing, technical, general")
    notes:    str | None     = None

# 2. Implementación con todas las capas de seguridad
class CreateTicketTool:
    name    = "create_ticket"
    timeout = 5  # segundos

    def execute(self, raw_params: dict) -> dict:
        # Validación — lanza ValidationError si algo está mal
        params = CreateTicketParams(**raw_params)

        # Dry-run check — ¿hay conflicto con un ticket abierto?
        existing = self.ticket_service.get_open(params.user_id)
        if existing and existing.subject.lower() == params.subject.lower():
            return {"warning": "duplicate_ticket", "existing_id": existing.id}

        # Ejecución con timeout
        with timeout(self.timeout):
            result = self.ticket_service.create(params)

        # Audit log — inmutable
        audit_log.record(tool=self.name, params=params.dict(),
                         result={"ticket_id": result.id}, user_id=params.user_id)

        return {"ticket_id": result.id, "status": "created"}

# 3. Registry — el agente solo conoce el registry, no las implementaciones
class ToolRegistry:
    def __init__(self):
        self._tools = {
            "create_ticket":     CreateTicketTool(),
            "get_order_status":  GetOrderStatusTool(),
            "send_notification": SendNotificationTool(),
            "schedule_callback": ScheduleCallbackTool(),
        }

    def execute(self, name: str, params: dict) -> dict:
        if name not in self._tools:
            raise ValueError(f"Herramienta desconocida: {name}")
        return self._tools[name].execute(params)

    def get_schemas(self) -> list[dict]:
        # Genera los schemas para el API de Anthropic automáticamente
        return [t.get_anthropic_schema() for t in self._tools.values()]

Resources

pydantic v2, redis-py, httpx (async)

Exercise 1

Implement the 4 tools with their tests

Each tool must have at least 3 tests: happy path, invalid parameters and timeout.

  1. Implement create_ticket with a mock of the ticket service
  2. Write a test: what happens if user_id is empty?
  3. Write a test: what happens if the service takes longer than 5 seconds?
  4. Implement get_order_status, send_notification and schedule_callback with the same structure
  5. Check that the ToolRegistry correctly generates the schemas for the Anthropic API

Module deliverable

Goes into the final project: ToolRegistry + SessionManager. 4 typed tools with validation, timeout and audit log. A centralized registry. A SessionManager for persistence between turns. The ToolRegistry is injected into the SupportAgent. When the agent decides to use a tool in its ReAct loop, it goes through the registry and never calls the service directly.

Module 8 · Phase 4 · Trade-offs & Debugging

Trade-offs & Optimization

Model routing, prompt caching, context compression and architecture decisions

The 4 trade-offs every senior engineer must master

Latency vs Quality

Haiku responds in <500ms. Sonnet takes 1-3s. Opus can take 5-10s. The question isn't "which one is better" but "which one does the user need in this context".

Cost vs Depth

Sonnet costs 12x more than Haiku. For classification (simple), Haiku is enough. For complex reasoning with tools, Sonnet is worth every cent.

Autonomy vs Control

More autonomy = a better user experience. More control = less risk of costly mistakes. The answer depends on how reversible the action is.

Agent vs Pipeline

If the flow always follows the same steps, a deterministic DAG is faster, cheaper and more predictable. The agent adds value when the input is ambiguous.

Senior judgment

The engineer who knows when NOT to use an agent is more valuable than the one who uses them everywhere. Asking "do I really need an agent here?" is the difference between elegant solutions and overly complex systems.

Model Routing: the highest-impact optimization

The principle: use the cheapest model that solves the case correctly. A lightweight classifier (Haiku) decides which model each query needs. 70-80% of support queries are simple and can be resolved with Haiku.

Cost reduction strategies

  • Prompt caching: the static part of the system prompt is cached. Anthropic offers a 90% discount on cached tokens. Always put the static part first.
  • Context compression: summarize long history instead of sending it in full. Saves 40-60% in long multi-turn conversations.
  • Batch API: a 50% discount for non-urgent tasks (evaluations, offline generation).
  • The 80/20 of cost: 80% of spend comes from the 20% longest requests. Optimize the tail, not the average.
class ModelRouter:
    def select(self, query: str, context: dict) -> ModelTier:
        # Regla 1: casos críticos siempre al modelo estándar
        if context.get("ticket_priority") == "critical":
            return ModelTier.STANDARD

        # Regla 2: clasificar complejidad con el modelo más barato posible
        complexity_prompt = f"""Clasifica esta consulta: '{query}'
Responde SOLO con: SIMPLE o COMPLEX
SIMPLE: saludos, estado de pedido, preguntas de FAQ
COMPLEX: problemas técnicos, disputas, múltiples pasos"""

        response = self.llm.call(
            [{"role": "user", "content": complexity_prompt}],
            model=ModelTier.FAST,  # Haiku para clasificar
            max_tokens=5
        )

        if "SIMPLE" in response.text:
            return ModelTier.FAST      # Haiku: 10x más barato
        return ModelTier.STANDARD       # Sonnet: balance ideal


class ContextCompressor:
    max_history_tokens = 3000

    def compress(self, history: list[dict]) -> list[dict]:
        if self._count_tokens(history) <= self.max_history_tokens:
            return history  # No necesita compresión

        # Mantener los últimos 3 turnos intactos (más relevantes)
        recent = history[-3:]
        older  = history[:-3]

        # Resumir los turnos más antiguos
        summary_prompt = f"Resume en 2 oraciones los puntos clave de esta conversación: {older}"
        summary = self.llm.call([{"role": "user", "content": summary_prompt}],
                                 model=ModelTier.FAST)

        return [{"role": "system",
                  "content": f"Contexto previo (resumido): {summary.text}"}] + recent

Resources

litellm, time.perf_counter, Anthropic prompt caching

Exercise 1

Model routing benchmark

Measure empirically how much model routing saves without sacrificing quality.

  1. Take 50 real (or simulated) support queries
  2. Run them all with Sonnet. Record the total cost and the correct resolution rate
  3. Implement the ModelRouter and run the same 50 queries
  4. Compare: how much did you save? Did the resolution rate drop?
  5. Tune the classifier's threshold until you get the best cost/quality balance

Module deliverable

Goes into the final project: ModelRouter + ContextCompressor + ADR. A working optimization module plus an ADR document with the system's measured trade-offs. The ModelRouter replaces the fixed model of the M1 LLMClient. The system now selects the model dynamically. The ContextCompressor kicks in automatically in the M5 MemoryManager.

Module 9 · Phase 4 · Trade-offs & Debugging

Probabilistic Debugging

Analysis framework, loop detection and failure reproducibility

Why debugging probabilistic systems is different

In deterministic systems, the same input produces the same output, always. In systems with LLMs, the same input can produce slightly different outputs on each call. That completely changes the debugging strategy.

The 3 most frequent types of failure

  • Hallucinations: the model generates incorrect information with high confidence. Cause: insufficient context or weak constraints in the prompt. Mitigation: RAG + explicit constraints + grounding checks.
  • Loops: the agent repeats the same action indefinitely. Cause: the tool fails but the model doesn't recognize it as an error. Mitigation: MAX_ITERATIONS + loop detector + circuit breaker.
  • Silent degradation: quality drops gradually with no visible alert. Cause: model drift on the provider side or prompt drift from accumulated edits. Mitigation: continuous evaluation + dashboard alerts.
Fundamental rule

An isolated failure is noise. A pattern of failures is a signal. Before changing the code, quantify: how many times does the same failure happen in 100 calls? If it's under 1%, document it and monitor. If it's over 5%, act.

A 5-step debugging framework

The right process

  • 1. Reproduce: store the full input (prompt, history, tool results, model, version). Without reproducibility, debugging is impossible.
  • 2. Isolate: does it fail in planning, execution or evaluation? Test each component separately with synthetic inputs.
  • 3. Trace: review the agent's <thinking>. Was the reasoning correct? Was the data right?
  • 4. Quantify: is it an isolated case or systemic? Run it 20+ times before drawing conclusions.
  • 5. Iterate: change ONE variable at a time. Without an A/B test there are no valid conclusions.
import hashlib, json
from datetime import datetime

class LoopDetector:
    def __init__(self, window: int = 3):
        self.window = window  # comparar los últimos N estados

    def check(self, history: list) -> bool:
        if len(history) < self.window:
            return False
        # Si los últimos N thoughts son iguales → loop detectado
        last_n = history[-self.window:]
        hashes = [hashlib.md5(json.dumps(h["thought"].tool_name).encode()).hexdigest()
                  for h in last_n]
        return len(set(hashes)) == 1  # todos iguales = loop

class FailureStore:
    def capture(self, context: dict, error: Exception, agent_history: list) -> str:
        failure_id = f"fail-{datetime.utcnow().strftime('%Y%m%d%H%M%S')}"
        record = {
            "id":             failure_id,
            "timestamp":      datetime.utcnow().isoformat(),
            "error_type":     type(error).__name__,
            "error_message":  str(error),
            "prompt_version": context.get("prompt_version"),
            "model":          context.get("model"),
            "full_context":   context,   # TODO: redactar PII antes de guardar
            "agent_history":  agent_history,
        }
        self.db.save(failure_id, json.dumps(record))
        return failure_id

    def replay(self, failure_id: str) -> dict:
        """Recupera el contexto completo para reproducir el fallo exactamente."""
        return json.loads(self.db.get(failure_id))

Resources

sqlite3, pytest fixtures, hashlib

Exercise 1

Analyze 2 real failures of the system

The best way to learn debugging is to analyze real failures, not simulated ones.

  1. Run the SupportAgent with 20 varied queries. The FailureStore captures everything that fails
  2. Pick the 2 most interesting failures from the store
  3. For each one: use the ReplayRunner to reproduce the failure exactly
  4. Inspect the agent's <thinking>: where did the reasoning go wrong?
  5. Write the analysis in docs/failure-analysis-report.md: root cause, proposed fix, regression test

Module deliverable

Goes into the final project: Debug Toolkit + Failure Analysis Report. FailureStore, LoopDetector, ReplayRunner and an analysis report of at least 2 real failures with root cause and proposed fix. The FailureStore connects to the Orchestrator. The LoopDetector wraps the SupportAgent's ReAct loop. Any unhandled exception is captured automatically with full context.

Module 10 · Phase 5 · Observability & Evaluation

LLMOps: Observability & Continuous Evaluation

Tracing, business metrics, prompt A/B testing and guardrails

You can't improve what you don't measure

The LLMOps module is the one that closes the loop. Without observability, the system is a black box that works (or doesn't) without anyone knowing why. With observability, every improvement decision is backed by data.

The metrics that matter, in order

  • Task completion rate: the most important metric. What % of conversations ended with the user's problem solved?
  • Escalation rate: % escalated to a human. If it rises → the agent is getting worse. If it drops a lot → it may be letting through cases it should escalate.
  • Cost per successful interaction: (tokens × price) / successful interactions. The system's efficiency metric.
  • p95 latency: the 95th percentile of latency, what 95% of users experience. The average lies.
  • Tool error rate: % of tool calls that fail. Points to problems with external APIs.
LLMOps principle

The metrics dashboard should be visible to the whole team, not just the technical side. A dashboard with business metrics + technical metrics in a single view eliminates 80% of prioritization debates.

LLM-as-judge: scalable automatic evaluation

Manually evaluating the quality of 1000 responses a week isn't feasible. The LLM-as-judge pattern uses a second LLM to evaluate the output of the first. The evaluator receives the original query, the generated response and the evaluation criteria.

Evaluator bias

LLM-as-judge has biases: it favors longer, more formal responses, or ones that sound "more confident". Always validate your judge against human evaluations on a sample before using it as the only source of truth.

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider

tracer = trace.get_tracer("nexus-support-agent")

class AgentTracer:
    def trace_request(self, trace_id: str, user_id: str):
        return tracer.start_as_current_span("agent_request",
            attributes={"trace_id": trace_id, "user_id": user_id})


class MetricsCollector:
    def record_interaction(self, result: AgentResult, context: dict):
        TASK_COMPLETION.inc(1 if result.success else 0)
        ESCALATION_RATE.inc(1 if result.escalated else 0)
        COST_COUNTER.inc(result.cost_usd)
        LATENCY_HISTOGRAM.observe(result.latency_ms)
        TOKEN_COUNTER.inc(result.total_tokens)


class LLMJudge:
    judge_prompt = """Evalúa esta respuesta de soporte.
Query: {query}
Respuesta: {response}

Puntúa 1-5 en cada criterio y responde SOLO JSON:
{{"relevance": 1-5, "accuracy": 1-5, "tone": 1-5, "completeness": 1-5,
  "overall": 1-5, "reasoning": "explicación breve"}}"""

    def evaluate(self, query: str, response: str) -> dict:
        result = self.llm.call(
            [{"role": "user", "content": self.judge_prompt.format(
                query=query, response=response)}],
            model=ModelTier.STANDARD  # el judge necesita buen criterio
        )
        return json.loads(result.text)


class EvalPipeline:
    def run(self, prompt_version: str) -> EvalReport:
        results = []
        for case in self.load_test_set():
            output = self.agent.run(case["input"], case["context"])
            score  = self.judge.evaluate(case["input"], output.answer)
            results.append({"case_id": case["id"], "score": score, "passed": score["overall"] >= 3})

        pass_rate = sum(1 for r in results if r["passed"]) / len(results)
        return EvalReport(results=results, pass_rate=pass_rate,
                           version=prompt_version, baseline=self.get_baseline())

Resources

opentelemetry, langsmith, prometheus, grafana, presidio (PII)

Exercise 1

Implement the complete metrics dashboard

The dashboard is the first thing you look at when something fails in production.

  1. Instrument the Orchestrator so every request produces the 5 defined metrics
  2. Spin up Prometheus + Grafana locally with Docker Compose
  3. Create a dashboard with task_completion_rate, escalation_rate, cost_per_interaction, p95_latency and tool_error_rate
  4. Run 50 simulated queries and check that the metrics update correctly
  5. Set up an alert: if escalation_rate rises more than 20% in 1 hour, alert the Slack channel
Exercise 2

Evaluation pipeline in CI

The test that keeps a prompt change from breaking the system in production.

  1. Create a GitHub Action that runs the EvalPipeline on every PR that modifies a file in prompts/
  2. The Action fails the PR if the pass_rate drops more than 5% against the baseline
  3. Make an intentionally bad prompt change and check that CI catches it
  4. Make a good change and check that CI approves it
  5. Document the process in the repository's README

Module deliverable

Goes into the final project: Observability Stack + Eval Pipeline in CI. Full instrumentation of the system: tracing, 5 business/technical metrics, LLM-as-judge, an automatic evaluation pipeline and safety guardrails. This module instruments every previous component. The eval pipeline connects to the M2 PromptLoader, the M3 CriticAgent and the M9 FailureStore, forming the complete continuous improvement loop.

Final project

Final project: Nexus Support Agent, an end-to-end multi-agent system

All the deliverables from the 10 modules integrated into an observable, optimized, production-ready customer support system.

Core

  • Automatic classification with a lightweight model
  • RAG over a knowledge base with reranking
  • Episodic memory per user
  • 4 typed external tools
  • Critic loop before responding

Safety

  • Human-in-the-loop with 5 criteria
  • 3-level cascading fallback
  • Circuit breaker per tool
  • PII detection on inputs/outputs
  • Global timeout per workflow

LLMOps

  • Full tracing with trace_id
  • Dashboard with 5 KPIs in real time
  • Automatic eval pipeline in CI
  • Dynamic model routing
  • Documented ADR backed by data

Repository structure

  • src/ agents/ llm/ memory/ tools/ safety/ observability/ optimization/ debug/
  • prompts/ support_agent/ classifier/ critic/ with versioning
  • evals/ test_set.jsonl · eval_pipeline.py · llm_judge.py
  • docs/ ADR-001.md · ADR-002.md · failure-analysis.md
  • tests/ unit/ integration/ coverage >70%
  • .github/ workflows/eval_on_pr.yml · ci.yml

Passing criteria

Criterion

Your progress is saved in this browser.