Strands harness: from one import to a production agent
A harness is everything around the model: tools, context, memory, permissions. This course has you run the Strands one, understand every default, and decide which to change before putting it to work.
- modules
- 8
- exercises
- 16
- final project
- 1
- Theory: Short, with one insight and one antipattern per module.
- Exercises: Two per module. You run them in your terminal and each has a result you can verify.
- Deliverable: One checklist per module. Your progress is saved in this browser.
- Final project: An agent with a session, memory, approval gate and durable state, with measurable criteria.
Based on the official Strands harness documentation (strandsagents.com/docs/user-guide/harness). It covers overview, quickstart, subagents, sessions and memory, context and caching, interventions, production and the configuration reference. Skills, MCP, background tasks and built-in tools are left for next steps. Model names and versions come from the docs as of October 2026: verify them before use.
Module 1 · Foundations
What a harness is and what it ships with
An agent is a model with tools inside a loop. The harness is everything else: the tuned prompt, context management, memory, the base tools. Strands ships it assembled behind a single import.
By the end: Run create_harness() with no arguments, list its defaults from memory and locate each one in the configuration reference.
One import, a ready agent
create_harness() (createHarness() in TypeScript) returns a standard Strands Agent. There is no wrapper or hidden abstraction: what changes are the defaults it comes assembled with.
What you get by default
- A model with reasoning on. Amazon Bedrock is the default provider.
- A tuned prompt: explore first, then act, confirm before anything irreversible, verify before finishing.
- Shell and file tools (read, write, edit) and web access.
- Context-window management and prompt caching.
- Long-term memory and sessions you resume with an id.
- A generalist subagent and a task checklist (todos).
- Agent Skills when present and programmatic tool calling.
Opinionated, not restrictive
Every default can be narrowed, swapped or turned off. As your case gets specific, you override what matters and keep the rest. If you want to build your own harness from scratch, the Strands Harness SDK takes the same configuration.
Configuration reference
| Option (Python) | TypeScript | Default | What it does |
|---|---|---|---|
model | model | bedrock/global.anthropic.claude-opus-5 | A provider/name string, a bare Bedrock id, or a Model instance. |
effort | effort | "auto" | Reasoning effort: auto, low, medium, high or off. Ignored for a Model instance. |
instructions | instructions | none | A domain block appended after the harness contract. Ignored when a full system prompt is passed. |
tools | tools | none | Your tools, added alongside the built-ins. |
plugins | plugins | none | SDK plugins, added alongside the built-in plugins. |
mcp_servers | mcpServers | none | MCP servers to connect: a JSON file path or a mapping. |
builtin_tools | builtinTools | shell, read, write, edit, web_fetch, web_search, programmatic_tool_caller, subagent | Built-in tools to enable, or [] for none. |
builtin_tools={"web_fetch": {"model": ...}} | builtinTools: { web_fetch: { model } } | provider default | The summarizer model web_fetch runs on. |
caching | caching | "auto" (on) | Prompt caching where the provider supports it. Off disables what the harness configures. |
context_manager | contextManager | "auto" | auto, agentic or off. Enables context management and offloading. |
session={"id": ...} | session: { id } | none | Persist and resume this conversation by id. |
session={"dir": ...} | session: { dir } | ./.agent/sessions | Where session state and offloaded artifacts live. |
skills | skills | ./.agent/skills | Directory (or list) scanned for Agent Skills. Off disables. |
builtin_plugins | builtinPlugins | ["todos", "environment"] | Built-in feature plugins to enable, or [] for none. |
memory | memory | on | File-based long-term memory. Off disables. |
memory={"dir": ...} | memory: { dir } | ./.agent/memory | Where the default memory store files live. |
memory={"stores": [...]} | memory: { stores } | none | Swap the memory backend while keeping the harness policy. |
interventions | interventions | none | Gate tool calls behind approval or a policy. |
background_tasks | backgroundTasks | { agentic: ['*'] } | Background-tasks policy. |
Because the result is a plain Agent, everything you know about the SDK (hooks, streaming, observability) still applies. The harness is a configuration factory, not a separate framework.
Treating it as a black box. The defaults write to disk (./.agent), can run shell commands, edit files and use Claude Opus 5 on Bedrock. Read them before putting it in front of real data.
First run with zero configuration
- Create an environment with Python 3.10 or newer and install strands-harness.
- Provide Bedrock credentials (
AWS_BEARER_TOKEN_BEDROCKor AWS credentials) and enable the model in the Bedrock console. - Run the code and wait for it to finish.
- Look at what appeared in the working directory.
# pip install strands-harness
from strands_harness import create_harness
agent = create_harness()
agent("Research the top three vector databases, compare pricing and limits, and write it up in comparison.md")import { createHarness } from '@strands-agents/harness'
const agent = await createHarness()
await agent.invoke("Research the top three vector databases, compare pricing and limits, and write it up in comparison.md") How to know it worked: comparison.md exists, written by the agent, along with a ./.agent/sessions folder. ./.agent/memory appears once memory extraction runs (every few turns). Without Bedrock, use model="anthropic/claude-sonnet-5" and continue with module 2.
Map the defaults against the reference
- Open the configuration reference table in this module.
- In a defaults.md file, write for each line of "what you get by default" which option controls it and how to turn it off.
- Check your table by building an agent with everything switchable turned off.
- Ask it to create a file and watch what it says.
from strands_harness import create_harness
agent = create_harness(
builtin_tools=[], # no shell, files, web, subagent
builtin_plugins=[], # no todos, no environment
memory=False,
session=False,
context_manager=False,
caching=False,
)
agent("Create a file named hello.txt with the word hi.")import { createHarness } from '@strands-agents/harness'
const agent = await createHarness({
builtinTools: [],
builtinPlugins: [],
memory: false,
contextManager: false,
caching: false,
// session off: see the "Persist sessions" page for your version
})
await agent.invoke('Create a file named hello.txt with the word hi.') How to know it worked: With no built-in tools the agent cannot create hello.txt: it only talks. Then switch one option back on at a time and see which enables what.
Recall
Answer out loud before opening the answer.
What does create_harness() return?
A standard Strands Agent, with no wrapper or hidden abstraction.
Name four things the agent has without you asking.
For example: shell and file tools, web access, long-term memory, context management, the generalist subagent and the todos checklist.
Module deliverable
Goes into the final project: Keep defaults.md. In the final project you decide which defaults to keep and which to change, and justify it in the README.
Module 2 · Foundations
Quickstart: CLI, library and model choice
There are three ways to start: ask your coding assistant to guide you, build the agent in the CLI, or use it as a library. All three end in the same agent.
By the end: Run the same agent from the CLI and from code, switch provider with one line and resume a conversation by id.
Three paths
- With your coding agent: paste the docs prompt into Codex, Claude Code or Kiro and it walks you through.
- CLI:
npm install -g @strands-agents/cli, run strands and choose Quickstart. The only thing you set up is the model provider. - Library:
pip install strands-harness(Python 3.10 or newer) or the@strands-agents/harnesspackage in TypeScript.
Choosing a model
The model is passed as provider/name. Amazon Bedrock is the default. The docs list Anthropic, OpenAI, Google and Ollama for running locally. In the CLI, strands --model anthropic/claude-sonnet-5 changes the model for that run only.
From the CLI to code
When you want to embed the agent, /export inside the CLI chat writes a Python or TypeScript project with your choices set on create_harness(...). The CLI is a ramp: build interactively, then move to code.
Sessions from day one
Sessions are on: each conversation is saved under ./.agent/sessions with a generated id. If you choose the id, a later run resumes the same conversation. In the CLI: strands --session-id api-design.
With Ollama the model runs on your machine and you need no cloud credentials: it is the cheapest way to iterate on configuration. A small model may fail more often at tool use, so try it on your real task before trusting it.
Pasting an API key into code. The CLI keeps a pasted key for the current session only; to reuse it, put it in your shell profile. In code, read it from the environment.
Same agent, three providers
- Put each provider key in the environment and run Ollama with ollama pull llama3.1.
- Run the script: same task, three models, one file per model.
- Note in one line which used tools best and which suits cheap iteration.
from strands_harness import create_harness
MODELS = ["anthropic/claude-sonnet-5", "openai/gpt-5.4", "ollama/llama3.1"]
for model in MODELS:
name = model.split("/")[0]
try:
agent = create_harness(model=model, session=False)
agent(
"Research the three most common strategies for versioning a REST API, "
f"compare their tradeoffs, and write a recommendation to api-versioning-{name}.md"
)
except Exception as err:
print(f"{model} failed: {err}")import { createHarness } from '@strands-agents/harness'
const models = ['anthropic/claude-sonnet-5', 'openai/gpt-5.4', 'ollama/llama3.1']
for (const model of models) {
const name = model.split('/')[0]
try {
const agent = await createHarness({ model })
await agent.invoke('Research the three most common strategies for versioning a REST API, compare their tradeoffs, and write a recommendation to api-versioning-' + name + '.md')
} catch (err) {
console.error(model + ' failed:', err)
}
} How to know it worked: You have one api-versioning-*.md file per provider that worked. If one failed, the message tells you which credential is missing.
CLI, export and a resumable session
- Install the CLI and run strands. Choose Quickstart, a provider and a model, then Save and Launch.
- Quit and enter again with an id you choose. Ask for a short task.
- Close, reopen with the same id and ask "what did I ask you before?".
- Inside the chat run
/exportand choose Python. Open the project and findcreate_harness(...).
npm install -g @strands-agents/cli
strands # setup, then chat
strands --model anthropic/claude-sonnet-5
strands --session-id curso-m2 # resume by id How to know it worked: On reopening with the same id the agent remembers the conversation. The exported project imports create_harness and carries your options.
Recall
Answer out loud before opening the answer.
What is the model argument format and what is the default?
provider/name, for example anthropic/claude-sonnet-5. The default is bedrock/global.anthropic.claude-opus-5.
What does /export do in the CLI?
It writes a Python or TypeScript project with your choices set on create_harness(...), exporting a ready-to-import agent.
Module deliverable
Goes into the final project: Pick the model and provider for the final project and note them in the README along with the environment variable it needs.
Module 3 · Configure the agent
Instructions, tools and subagents
The harness ships a tuned prompt and base tools. Your job is to add your domain context, your own tools and, when needed, specialists to delegate to.
By the end: Add instructions, add a specialist with as_tool() and use the generalist subagent to keep the main context clean.
instructions does not replace the prompt
The instructions parameter appends a domain block after the harness contract (explore first, confirm before anything irreversible, verify before finishing). If you pass a full system prompt, instructions is ignored and you lose that contract.
Your tools live alongside the built-ins
tools adds yours to shell, read, write, edit, web_fetch, web_search, programmatic_tool_caller and subagent. builtin_tools chooses which built-ins stay: an empty list leaves none.
Subagents: delegate to protect context
A subagent is an agent the main agent calls like a tool. Its intermediate work stays out of the main conversation: only the final answer comes back.
- Your specialists: an Agent with name, description and
system_prompt, passed withas_tool(). Each call starts fresh, with no accumulated state, and the name must be unique across tools. - generalist: comes enabled. It inherits model, reasoning, caching, context management, built-in tools, plugins, subagents, interventions and sandbox, but runs a generic role prompt instead of your instructions. It starts blank: the call must carry everything it needs.
When to delegate
When a subtask would flood the context: searching many files, a multi-step change, or open-ended exploration where you only need the conclusion. The generalist runs in the background, so the main agent can keep working.
A subagent inherits interventions and the sandbox. By design it cannot become a back door around the approval you set on the main agent.
A specialist with a vague description. The model decides when to call a tool by its name and description: "helper" says nothing. Write what it does, what it takes and what it returns.
A specialist with a name and description
- Define a researcher Agent with name, description and
system_prompt. - Pass it with
as_tool()and add an instructions block saying when to delegate. - Ask for a task that needs research and writing and watch for the researcher call in the output.
- Experiment: change the description to "helper" and repeat.
from strands import Agent
from strands_harness import create_harness
researcher = Agent(
name="researcher",
description="Researches a topic and returns a concise, sourced summary.",
system_prompt="Research the given topic and return a concise, sourced summary.",
)
agent = create_harness(
instructions=(
"You write short technical briefs. "
"Delegate research to the researcher tool, then write the brief to brief.md."
),
tools=[researcher.as_tool()],
)
agent("Brief me on prompt caching in LLM APIs, in under 300 words.")import { Agent } from '@strands-agents/sdk'
import { createHarness } from '@strands-agents/harness'
const researcher = new Agent({
name: 'researcher',
description: 'Researches a topic and returns a concise, sourced summary.',
systemPrompt: 'Research the given topic and return a concise, sourced summary.',
})
const agent = await createHarness({
instructions: 'You write short technical briefs. Delegate research to the researcher tool, then write the brief to brief.md.',
tools: [researcher.asTool()],
})
await agent.invoke('Brief me on prompt caching in LLM APIs, in under 300 words.') How to know it worked: brief.md was created and you saw the researcher tool call. With the vague description, note whether the agent stops delegating or delegates worse.
Generalist: delegate to keep context clean
- Put a
./reposfolder with several README.md files (your repos or copies). - Run the script: the same task with the generalist on and with
builtin_tools={"subagent": False}. Each run uses its own session directory. - Compare the size of each session directory and note what you saw in the output.
from strands_harness import create_harness
TASK = "Read every README.md under ./repos, then write overview.md with one line per project."
with_delegate = create_harness(
session={"id": "m3-with", "dir": "./.s-with"},
)
without_delegate = create_harness(
session={"id": "m3-without", "dir": "./.s-without"},
builtin_tools={"subagent": False},
)
with_delegate(TASK)
without_delegate(TASK)
# then, in the terminal: du -sh ./.s-with ./.s-without How to know it worked: There are two overview.md files and two session directories. It is expected that the run without the delegate leaves more content in the main conversation; if you do not see it, note why (for example, the offloader moved bulky results).
Recall
Answer out loud before opening the answer.
What happens to instructions if you pass system_prompt?
It is ignored: system_prompt replaces the prompt the harness builds and you lose its contract.
What does the generalist inherit and what does it not?
It inherits model, reasoning, caching, context management, tools, plugins, subagents, interventions and sandbox. It does not inherit your instructions (it uses a generic role prompt) nor the conversation (it starts blank).
Module deliverable
Goes into the final project: The final project carries a specialist of your own (for example a source checker) and an instructions block with your domain rules.
Module 4 · Configure the agent
State: sessions, checkpoints and memory
An agent can remember in two different ways and they are not the same. A session resumes an exact conversation. Memory carries durable facts across conversations. Mixing them up gives agents that forget what they should not or drag along what does not belong.
By the end: Decide between a session, memory or both, and prove it with a run that survives a restart and one that leaves no trace.
Two different questions
A session (checkpoint) persists one conversation so you can resume that exact task after a restart. Long-term memory distills durable facts and recalls them in any conversation, with or without a session. They are orthogonal: you can run with neither, either or both. Both are on by default, but you only resume a conversation if you give a session id.
- Resume a task after a restart: a session, with
session={"id": ...}. - Carry facts, preferences or decisions across unrelated runs: memory, which is already on.
- Both: a session id and memory left on.
- A one-shot task that leaves no trace: session=False and memory=False.
How memory works
The harness distills durable facts into files under ./.agent/memory, searches them before each turn and folds the top matches into context. The agent also gets a search_memory tool for on-demand recall. Extraction runs in the background every few turns on a small model, so keeping memory costs little.
Backends: what ships and what does not
LocalFileStorage(default, atomic writes),S3Storage(production and multi-instance) andInMemoryStorage(tests).- A custom backend implements four async methods: write, read, delete and list.
- SQLite and PostgreSQL are not first-party: you write them against Storage or
MemoryStore. Redis or Valkey exist only through the communitystrands-valkey-session-managerpackage (Python).
Concurrency, isolation and deletion
- One writer per conversation. The managers take no distributed lock: two invocations on the same id overwrite each other and neither errors (last write wins). If you fan out, add your own lock or route each id to a single worker.
- Isolation by namespace and scoped stores: memory supports a store per tenant rather than one shared store.
- Deletion: deleting a session removes its root directory (or the prefix in S3, which needs
s3:DeleteObject). Memory is plain files: removing a tenant memory means deleting its directory or store. - The session directory is a trusted data store: restrict its permissions to the agent process. The SDK does not block symlinks inside it.
A session is for continuing a task; memory is for knowledge that outlives the task. If you ask yourself "do I want this tomorrow in another conversation?", the answer picks which one to use.
Two workers on the same session id. There is no error: the second overwrites the first one turns and you find out once the conversation is already broken.
Resume after a restart
- Run the first block with the id m4-api and let it finish.
- Stop the process entirely. That is the restart.
- In a new process run the second block with the same id.
- Look at
./.agent/sessionsto see where the state landed.
# run 1
from strands_harness import create_harness
agent = create_harness(session={"id": "m4-api"})
agent("List three strategies for versioning a REST API.")
# run 2, in a brand new process
agent = create_harness(session={"id": "m4-api"})
agent("Which of those would you pick for an API with external customers, and why?")// run 1
import { createHarness } from '@strands-agents/harness'
const agent = await createHarness({ session: { id: 'm4-api' } })
await agent.invoke('List three strategies for versioning a REST API.')
// run 2, in a brand new process
const again = await createHarness({ session: { id: 'm4-api' } })
await again.invoke('Which of those would you pick for an API with external customers, and why?') How to know it worked: The second answer refers to the three strategies without you repeating them.
Memory across unrelated conversations
- In conversation A, ask it to remember a fact and take three or four more turns to give extraction time.
- In conversation B (another session id) ask about that fact.
- In conversation C, with memory=False, ask the same question.
- Look at the files in
./.agent/memory.
from strands_harness import create_harness
a = create_harness(session={"id": "m4-a"})
for turn in [
"Remember: my stack is FastAPI and Postgres, and I prefer answers as short tables.",
"Give me one tip about queues.",
"Give me one tip about caching.",
"Give me one tip about retries.",
]:
a(turn)
b = create_harness(session={"id": "m4-b"}) # new conversation
b("What stack do I use?")
c = create_harness(session={"id": "m4-c"}, memory=False) # memory off
c("What stack do I use?") How to know it worked: B answers with the stack and C does not. If B does not know, background extraction may not have run yet: add turns in A and look at ./.agent/memory.
Recall
Answer out loud before opening the answer.
What is the difference between a session and memory?
A session persists one conversation so you can resume it by id. Memory distills durable facts and recalls them in any conversation.
What happens if two processes write to the same session?
They overwrite each other: no distributed lock, last write wins, no error. Use one writer per id.
Module deliverable
Goes into the final project: The final project uses one session id per task and leaves memory on with its directory on durable storage. Note in the README who writes each id.
Module 5 · Configure the agent
Context and caching
A model only reads a limited span of text at once. The harness keeps the conversation inside that limit and reuses what does not change between turns, so each step is cheaper and faster.
By the end: Understand what context_manager and caching do, change their mode and know where bulky results end up.
Context management
With context_manager on, the harness summarizes older turns as the conversation grows and appends an offloader: bulky tool results move to storage and are replaced by a short preview and a reference the agent can follow to pull the full content back when it truly needs it.
- auto (default) and agentic select the SDK context strategy. Both keep the offloader on.
- Turning it off (False or null, or off in the CLI) also turns off offloading: the full history stays in the window and you own its size.
- With an active session, offloaded artifacts persist under the session directory. Without one they go to a temporary directory that does not outlive the process.
Prompt caching
It reuses what does not change between turns (system prompt, tool definitions and prior conversation), so the stable prefix of a long conversation is cheaper and faster to process. It is on by default.
- On Amazon Bedrock and Anthropic direct, the harness configures cache points and cached tool definitions.
- On OpenAI, Google and bedrock-mantle caching happens automatically server-side: there is nothing to configure.
- Turning caching off has no effect where it is automatic. Enabling it explicitly on a pre-built Model instance is ignored with a warning: configure it on the instance itself.
The offloader changes what the model sees, not what is lost: the full content stays stored and the agent can pull it back. That is why an active session is useful if you want to audit those artifacts later.
Turning context_manager off "to see everything" on a long task. Without it the full history stays in the window and the limit is yours to manage: the task can hit that limit.
What the model sees and what stays stored
- Run a task that produces bulky results (long pages), once with managed context and once with
context_manager=False. Each run uses its own session directory. - Look in each directory for a context subdirectory: the docs say the context manager stash lives under context/.
- Note what it holds and how much each directory weighs.
from strands_harness import create_harness
TASK = "Fetch three long documentation pages about HTTP caching and write a comparison to caching.md"
managed = create_harness(session={"id": "m5-on", "dir": "./.s-on"})
unmanaged = create_harness(session={"id": "m5-off", "dir": "./.s-off"}, context_manager=False)
managed(TASK)
unmanaged(TASK)
# terminal: find ./.s-on ./.s-off -type d -name 'context*' ; du -sh ./.s-on ./.s-off How to know it worked: In ./.s-on there should be a context directory with offloaded artifacts and in ./.s-off there should not. If your task did not produce results large enough, there will be no offloading: try longer pages.
Caching: what the harness sets up and what it does not
- For each provider you tried in module 2, note whether caching is set up by the harness (Bedrock, Anthropic) or automatic (OpenAI, Google, bedrock-mantle).
- Run the same multi-turn conversation with default caching and with caching=False and compare the total time.
- If your provider reports cached tokens, note that figure.
import time
from strands_harness import create_harness
TURNS = [
"Explain HTTP caching in 3 bullets.",
"Now ETag versus Last-Modified.",
"Now the Cache-Control directives.",
"Summarize our whole chat in 2 lines.",
]
for label, extra in [("caching auto", {}), ("caching off", {"caching": False})]:
agent = create_harness(session=False, **extra)
start = time.time()
for turn in TURNS:
agent(turn)
print(label, round(time.time() - start, 1), "s") How to know it worked: You have two timings and a note per provider. Do not expect a fixed difference: it depends on the provider and the prefix length. What matters is knowing who sets up caching in your case.
Recall
Answer out loud before opening the answer.
What does the offloader do?
It moves bulky tool results to storage and replaces them with a short preview and a reference to pull them back.
What happens with context_manager=False?
Offloading is also turned off: the full history stays in the window and you manage its size.
Module deliverable
Goes into the final project: The final project leaves context_manager on auto, the session on a durable directory (the artifacts live there) and documents in the README whether caching is set by the harness or the provider.
Module 6 · Control and production
Interventions: put a gate on tool calls
By default every tool call runs. An intervention decides whether a call runs, and it is the documented way to put human approval or a policy in front of the shell and file edits.
By the end: Gate calls with the ask and smart presets, a natural-language rule and a Cedar policy, and check that the subagent does not bypass the gate.
Four ways to gate
"ask": asks for approval on every call."smart": uses the SDK risk classifier and gates only risky calls.- A string that is not a preset and does not end in .cedar is a natural-language risk policy: it becomes the classifier prompt, like smart with your own rubric.
- A string ending in .cedar loads a Cedar policy (needs
strands-agents[cedar]in Python or@cedar-policy/cedar-wasmin TypeScript). Inline Cedar text is not auto-detected: pass aCedarAuthorizationinstance.
Layers and custom handlers
For what the presets do not cover, build the SDK handler and pass it: an instance passes through untouched. You can also pass a list to layer, for example a Cedar policy with a human-approval gate. An agent registers at most one handler per name, so layer different kinds, not duplicates.
Pause and resume
When a gate asks for approval, execution is interrupted. With the SDK Agent the pattern is: if result.stop_reason is "interrupt", respond with interruptResponse and the interrupt id to resume. Because the harness returns a standard Agent, confirm in your version how it is exposed before building a UI on top.
The generalist inherits the policy
The built-in subagent inherits whatever you set in interventions, so it cannot bypass the main agent gate.
A natural-language rule is flexible, but a model evaluates it: it is for catching the doubtful, not for guaranteeing the impossible. What must never happen belongs in a Cedar policy or a sandbox, which are deterministic.
Leaving interventions unconfigured on an agent with shell and file editing that receives third-party input. The default applies none: every call proceeds.
Presets and resuming
- Create an agent with interventions=
"ask"and ask it to create and delete a file. - When it interrupts, approve the first call and deny the delete.
- Repeat with
"smart"and note which calls it gated and which it let through.
from strands_harness import create_harness
agent = create_harness(interventions="ask") # then try "smart"
result = agent("Create temp.txt with the word hi, then delete it.")
# pattern from the SDK human-in-the-loop docs: verify it on your version
while result.stop_reason == "interrupt":
interrupt = result.interrupts[0]
print("Approval needed:", interrupt)
answer = input("approve? (yes/no) ")
result = agent([{"interruptResponse": {"interruptId": interrupt.id, "response": answer}}]) How to know it worked: With ask, every tool asks for approval. If you answer "no" to the delete, temp.txt still exists. With smart note what you observe: the classifier is a model and may decide differently each run.
Natural-language rule and inherited gate
- Create an agent with the rule "Ask before deleting files or making any network request."
- Ask for something that needs the web and check that
result.stop_reasonis"interrupt". - Ask for a task that delegates to the generalist (read many files, then delete one) and check that the delete is gated too.
- Write down one action you would always block and why it deserves Cedar or a sandbox instead of natural language.
from strands_harness import create_harness
agent = create_harness(
interventions="Ask before deleting files or making any network request.",
)
result = agent("Search the web for the latest Python release and write it to python.txt")
print(result.stop_reason) # expect "interrupt" when a network call is gatedimport { createHarness } from '@strands-agents/harness'
const agent = await createHarness({
interventions: 'Ask before deleting files or making any network request.',
})
const result = await agent.invoke('Search the web for the latest Python release and write it to python.txt')
console.log(result.stopReason) // expect "interrupt" when a network call is gated How to know it worked: stop_reason is "interrupt" when a network call is attempted. In step 3 the delete inside the delegation also asks for approval, because the generalist inherits the policy.
Recall
Answer out loud before opening the answer.
What is the default for interventions?
None: every tool call runs.
When a natural-language rule and when Cedar?
A natural-language rule is evaluated by a model: it is for escalating the doubtful. Cedar is a programmatic policy: it is for what must never happen.
Module deliverable
Goes into the final project: The final project carries a gate: a natural-language rule for the doubtful and, if your environment allows, a .cedar for the forbidden. The measurable criterion is that a denied destructive action does not run.
Module 7 · Control and production
Production: what it writes, what it runs and where state lives
The harness returns a normal Agent, so deploying, observing and securing it follows the SDK guides. What is specific to the harness are its defaults: they write to disk and can run code.
By the end: Inventory the risks of the defaults, move state to durable storage and package the agent in a container.
Deploy: same as any Agent
Every target in the SDK deploy guides (Lambda, Fargate, EKS, Bedrock AgentCore, Docker and more) works unchanged: where the guide builds Agent(), you build create_harness().
State on ephemeral storage
Sessions and memory write by default to directories under ./.agent. On a container or serverless target, point session={"dir": ...} and memory={"dir": ...} at durable storage (a mounted volume) or a backend of your own, via memory={"stores": [...]} or a session manager, so state outlives a single instance.
Observe
It is a standard Agent: SDK telemetry (traces, metrics and logs) works unchanged and there is nothing harness-specific to wire up. To measure quality, the Evals SDK tests and scores harness runs like any other agent.
Two defaults that need a security decision
programmatic_tool_callerruns model-authored code in Monty: it isolates the code, but not the tools that code calls. If the agent handles untrusted input, run it inside an SDK sandbox or drop the tool.- The default agent can run shell commands and edit files: gate what it does with interventions and confine what it can reach with a sandbox.
- The SDK safety guides (guardrails, PII redaction, trusted message history) apply directly.
An allowlist is easier to audit than a blocklist: instead of asking what to turn off, you declare with builtin_tools exactly which tools exist.
Deploying in a container without mounting the session and memory directories. Everything works until the first redeploy: then the agent loses its conversation and what it learned.
Risk surface inventory
- For each built-in tool (shell, read, write, edit,
web_fetch,web_search,programmatic_tool_caller, subagent) write what harm it could do with untrusted input. - Decide which stay and build the agent with that list in
builtin_tools. - Ask for something outside the list and check that it cannot do it.
from strands_harness import create_harness
# example allowlist for a research agent that must not touch the filesystem
agent = create_harness(
builtin_tools=["web_search", "web_fetch"],
builtin_plugins=["todos"],
interventions="smart",
)
agent("Write a file named x.txt with the word hi.") # it should be unable to How to know it worked: The agent says it cannot write files. Your risk table explains why each tool you kept is necessary.
Container with durable state
- Create the three files: agent.py, Dockerfile and compose.yaml. State lives in the /data volume.
- Run a task with a fixed
SESSION_ID. - Destroy the container with
docker compose down. The volume is kept. - Run another task with the same
SESSION_IDand check that it resumes.
# agent.py
import os
import sys
from strands_harness import create_harness
agent = create_harness(
model=os.environ.get("HARNESS_MODEL", "anthropic/claude-sonnet-5"),
session={"id": os.environ["SESSION_ID"], "dir": "/data/sessions"},
memory={"dir": "/data/memory"},
)
agent(sys.argv[1])FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir strands-harness
COPY agent.py .
ENTRYPOINT ["python", "agent.py"]services:
agent:
build: .
environment:
- SESSION_ID=${SESSION_ID:-demo}
- HARNESS_MODEL=${HARNESS_MODEL:-anthropic/claude-sonnet-5}
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} # use the variable your provider expects
volumes:
- agent-data:/data
volumes:
agent-data:docker compose run --rm agent "List three API versioning strategies."
docker compose down # the volume stays
docker compose run --rm agent "Which of those would you pick, and why?"
# to prove the volume matters: docker compose down -v and repeat How to know it worked: The second run resumes the conversation even though the original container is gone. With docker compose down -v (which deletes the volume) it starts from scratch: that proves state lives in the volume.
Recall
Answer out loud before opening the answer.
Why move session and memory out of ./.agent in a container?
Because by default they write to ./.agent, and on a container or serverless target that is ephemeral storage: state is lost with the instance.
What does Monty isolate and what does it not?
It isolates the model-authored code, but not the tools that code calls.
Module deliverable
Goes into the final project: The final project ships as a compose.yaml with a volume for session and memory, a tool allowlist and a README with the risks you accepted.
Module 8 · Beyond the defaults
Leaving the defaults: dropping down to the SDK
The harness is an opinionated factory. When an option is not enough, any argument the factory does not name goes through to the Agent constructor, and you can drop down to the SDK without losing the configuration you already have.
By the end: Know what a passthrough argument overrides, which pieces the harness exports and when it makes sense to build your own harness with the SDK.
Passthrough to the Agent
Any argument the factory does not name is forwarded to the Agent constructor. An explicit value passed this way wins over the default it corresponds to: passing system_prompt overrides the prompt the harness builds from instructions, and passing memory_manager overrides memory. In TypeScript, HarnessAgentOptions extends the SDK AgentConfig (minus the fields the harness manages), so other fields such as retryStrategy pass through verbatim.
Exported pieces
Beyond the factory, the harness exports what it composes: HARNESS_CONTRACT, build_system_prompt (buildSystemPrompt in TypeScript), resolve_memory and resolve_interventions (Python only). They let you inspect or recompose instead of rewriting from scratch.
Which override level to use
- One parameter changes: a
create_harnessargument. - The whole prompt changes:
system_prompt, knowing you lose the contract and instructions is ignored. - The memory backend changes:
memory={"stores": [...]}keeps the harness policy;memory_managerreplaces it. - The session backend changes: pass your own session manager, which wins over the one the harness builds from session.
- Everything changes: the Strands Harness SDK directly, with the same configuration you already know.
Pick the lowest level that solves your problem. memory={"stores": [...]} changes where data is kept and preserves the harness policy; memory_manager replaces the entire policy.
Passing system_prompt to "add a rule". To add domain context there is instructions: with system_prompt you replace everything and lose the explore, confirm and verify contract.
instructions versus system_prompt
- Create a notes.txt file in the working folder.
- Build two agents with the same rule: one with instructions and one with
system_prompt. - Ask each to delete notes.txt (restore the file between runs) and note whether the agent confirms before acting.
from strands_harness import create_harness
RULE = "Always answer in exactly two sentences."
a = create_harness(instructions=RULE, session=False) # keeps the harness contract
b = create_harness(system_prompt=RULE, session=False) # replaces it entirely
for agent in (a, b):
agent("Delete notes.txt from the current folder.") How to know it worked: A follows the harness contract (confirm before anything irreversible); B no longer has it. If both behave the same, note it: the model may confirm on its own, but with system_prompt there is no contract guarantee.
Read the contract and change a store
- Import
HARNESS_CONTRACTand print it. Underline the rules you recognize from your defaults.md. - Change only the memory directory with
memory={"dir": ...}. - After a few turns, check that memory files show up in the new directory and not in
./.agent/memory. - Write down when you would use stores instead of dir.
# if the import path differs in your version, see "Compose with the Strands Harness SDK"
from strands_harness import create_harness, HARNESS_CONTRACT
print(HARNESS_CONTRACT)
agent = create_harness(memory={"dir": "./my-memory"}, session={"id": "m8"})
for turn in [
"Remember that my deploy target is a single VPS with Docker Compose.",
"Give me one tip about volumes.",
"Give me one tip about backups.",
]:
agent(turn)
# terminal: ls ./my-memory How to know it worked: You could read the contract the harness adds and saw ./my-memory fill up instead of ./.agent/memory (it can take a few turns because of background extraction).
Recall
Answer out loud before opening the answer.
Which wins: an argument passed to the Agent or the harness default?
The explicit argument. Passing system_prompt overrides the prompt built from instructions, and passing memory_manager overrides memory.
When does it make sense to drop to the SDK?
When several structural pieces change or you want your own harness from scratch. If only one piece changes, passthrough is enough.
Module deliverable
Goes into the final project: In the final project README, an "Overrides" section lists every parameter you changed from the default and why.
Final project
A research agent with guardrails
You build an agent that researches, writes a report and does so auditably: it resumes its conversation, remembers what matters, asks permission for anything destructive and keeps its state when the container dies.
What you deliver
- A repository with compose.yaml, Dockerfile and agent.py, runnable with a single docker compose run command.
create_harnesswith: explicit model, domain instructions, a specialist viaas_tool(),builtin_toolsas an allowlist, interventions, session with one id per task and memory, both with directories on the volume.- A README gathering what you wrote in each module: defaults kept and changed (1), model and environment variable (2), specialist and rules (3), who writes each id (4), who sets up caching (5), what goes to Cedar or a sandbox (6), accepted risks (7) and overrides (8).
Suggested order of work
- Skeleton with the chosen model and a research task that runs end to end (modules 1 and 2).
- Domain instructions and the specialist (module 3).
- State on the volume: a session per task and memory (modules 4, 5 and 7).
- Approval gate and tool allowlist (modules 6 and 7).
- README with the decisions and overrides (modules 1 and 8).
- Verification of the six criteria below from a clean clone.
Passing criteria
| Criterion | How to measure | Module |
|---|---|---|
Run a task, run docker compose down, repeat with the same SESSION_ID and the answer uses the earlier context. | 4, 7 | |
A fact stated in task A shows up in the answer to task B, which uses another SESSION_ID. | 4 | |
| A destructive action you deny does not run: the file still exists. | 6 | |
| A destructive action requested through the specialist or the generalist is also gated. | 3, 6 | |
builtin_tools is an explicit list and each tool has a justification line in the README. | 1, 7 | |
| From a clean clone, with the provider key, the README is enough to run the five points above. | 8 |
Your progress is saved in this browser.
Next steps
These doc pages are outside the course. Each one completes a piece you have already seen from the outside.
- Load agent skills: The
./.agent/skillsdirectory the harness scans by default. - Connect MCP servers: The
mcp_serversoption: a JSON path or a mapping. - Run work in the background: The
background_taskspolicy and how the generalist runs in the background. - Shell and file tools: The shell, read, write and edit tools in detail.
- Web access:
web_fetch,web_searchand the summarizer model. - Programmatic tool calling: The tool that runs model code in Monty.
- Task tracking and environment: The built-in todos and environment plugins.
- Compose with the Strands Harness SDK: How the harness sits on the SDK and how to go beyond it.
- Versioning & Support: Versioning policy: read it before pinning dependencies.