I built a small agent to watch this happen: three Python files, no agent framework, every step sent to Splunk before the next one runs. One question produced 25 events. This guide starts from zero, shows the code behind each part, then reads those 25 events one by one.

What are the parts of an AI agent?

Four parts, three of them on your machine. The model sits behind an API.

Part File What it does What it never does
Model (LLM) remote, through OpenRouter Reads the context, reasons, returns text or a tool call as JSON Runs code, reads files, opens connections
Agent loop agent.py, 151 lines Builds the context, calls the model, runs the requested tools, stops at a final answer or after 6 turns Runs a tool outside the MCP server
MCP server mcp_server.py, 60 lines Holds the tools and their descriptions, executes them inside a sandbox Talks to the model
Logger hec_logger.py, 67 lines Sends every step to Splunk and stops the agent if Splunk does not acknowledge Lets the model write into the logs

The only part that reasons is the model. Everything that acts is ordinary code you can read, test and log.

Architecture of the AI agent lab: the agent loop and the MCP server on one VM, Splunk on another, NAT egress to the model API, hypervisor and private ranges blocked

How does the agent start the MCP server?

As a child process. The agent launches mcp_server.py with the same Python interpreter, then talks to it through the child’s standard input and output. No network port is opened.

server = StdioServerParameters(command=sys.executable,
                               args=[str(HERE / "mcp_server.py")],
                               env=dict(os.environ))
async with stdio_client(server) as (read, write):
    async with ClientSession(read, write) as mcp:
        init = await mcp.initialize()               # handshake
        tools = (await mcp.list_tools()).tools      # discovery

The code is asynchronous because the client reads the server’s replies while it keeps writing requests on the other stream. stdio_client starts the subprocess, and ClientSession speaks the protocol.

The protocol is JSON-RPC 2.0, one JSON object per line. The MCP specification calls this the stdio transport: newline-delimited messages over the standard streams of a subprocess the client launched. The first exchange is a handshake:

{"jsonrpc": "2.0", "id": 1, "method": "initialize",
 "params": {"protocolVersion": "2025-11-25", "capabilities": {...},
            "clientInfo": {...}}}

In my session, the server answered with its name, lab-files, protocol version 2025-11-25, and its capabilities: tools, prompts and resources, none of them able to change during the session (listChanged: false). Then the agent asked for the tool list with tools/list.

What is inside an MCP server?

A list of functions with a description each. In the MCP Python SDK, FastMCP turns a decorated Python function into a tool:

mcp = FastMCP("lab-files")

@mcp.tool()
def read_file(path: str) -> str:
    """Read one text file from the lab documents folder. Use a path returned by list_files."""
    target = _safe_path(path)          # resolve the path, refuse anything outside the sandbox
    raw = target.read_bytes()
    return raw[:20_000].decode("utf-8", errors="replace")

Three things happen in that decorator:

  • The function name becomes the tool name: read_file.
  • The docstring becomes the tool description, the text the model reads to decide when to use it.
  • The type hints become the input schema: one required string, path.

This is the entry the agent received for read_file in tools/list, copied from the trace:

{"name": "read_file",
 "description": "Read one text file from the lab documents folder. Use a path returned by list_files.",
 "inputSchema": {"type": "object", "title": "read_fileArguments",
                 "properties": {"path": {"title": "Path", "type": "string"}},
                 "required": ["path"]}}

Only decorated functions exist for the model. _safe_path sits in the same file, but no request can call it directly. This server exposes two tools, list_files and read_file, and nothing that writes, sends or deletes.

The description deserves a second look. It is plain text written by whoever wrote the server, and the model treats it as guidance. A server update can change what the model is told to do without a single change to agent.py, because the agent discovers tools at runtime. That is the entry point of tool poisoning, so the agent logs a SHA-256 of every description at the start of each session.

How does the model decide to call a tool?

The agent converts the MCP tools into the format the model API expects, then sends them with every request:

openai_tools = [{"type": "function",
                 "function": {"name": t.name,
                              "description": t.description,
                              "parameters": t.inputSchema}}
                for t in tools]

resp = llm.chat.completions.create(model=MODEL, messages=messages, tools=openai_tools)

The model does not run read_file. It returns a message that asks for it, with the arguments serialized as a JSON string, in the format OpenRouter documents for tool calling. This is the model’s reply at turn 2 of my session, from the trace:

{"content": null,
 "tool_calls": [
   {"id": "call_LvjcsHyxoSJyHkZQioVFmaHd", "name": "read_file",
    "arguments": "{\"path\":\"incident-report-2026-09.md\"}"},
   {"id": "call_SguvWH6kAl0H6UeetDgmoVLG", "name": "read_file",
    "arguments": "{\"path\":\"access-policy.md\"}"}],
 "finish_reason": "tool_calls"}

No text, two requests, and a finish reason of tool_calls. The reasoning happens inside the model, and what comes out is a request. Whether the request runs is the loop’s decision.

What does the agent loop look like in code?

A for loop with a turn limit. Here it is without the logging lines:

messages = [{"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": question}]

for step in range(1, MAX_STEPS + 1):                       # MAX_STEPS = 6
    resp = llm.chat.completions.create(model=MODEL, messages=messages, tools=openai_tools)
    msg = resp.choices[0].message
    if not msg.tool_calls:                                  # no tool requested
        answer = msg.content                                # final answer
        break
    messages.append({"role": "assistant", "content": msg.content or "",
                     "tool_calls": [...]})                  # keep the request in the history
    for call in msg.tool_calls:
        args = json.loads(call.function.arguments or "{}")
        result = await mcp.call_tool(call.function.name, args)   # JSON-RPC tools/call
        text = "\n".join(c.text for c in result.content)
        messages.append({"role": "tool", "tool_call_id": call.id, "content": text})

Every line maps to a decision you can secure:

Line Decision Control in the lab
range(1, MAX_STEPS + 1) How long the agent may run 6 turns, then max_steps_reached
chat.completions.create(...) What the model sees Full context logged in trace mode
if not msg.tool_calls When the agent stops Final answer logged before it is returned
mcp.call_tool(...) What actually runs Only the server’s two read-only tools exist
messages.append({"role": "tool", ...}) What goes back into the context Logged before it reaches the model

Why does the context grow at every turn?

Because the model keeps no memory between calls. Each request carries the full history, so the loop appends the model’s request and every tool result before asking again.

Turn Messages sent Prompt tokens What was added
1 2 136 System prompt and question
2 4 173 The model’s call to list_files and its result
3 7 446 The model’s two calls to read_file and the two documents

The prompt more than tripled between turn 1 and turn 3, and most of the jump comes from the 985 bytes of the two documents. A long tool result is paid for again at every following turn, one reason the server caps each read at 20,000 bytes. A document read at turn 2 also stays in the context for the rest of the session, so any instruction hidden inside it keeps reaching the model.

How is every step logged before the agent acts?

Through one method that refuses to fail silently:

def send(self, sourcetype: str, event: dict) -> None:
    payload = {"time": time.time(), "sourcetype": sourcetype, "index": self.index, "event": event}
    try:
        r = self.http.post(self.url, json=payload, timeout=5, verify=self.verify)
    except requests.RequestException as exc:
        raise LoggingError(f"Splunk HEC unreachable: {exc}") from exc
    if r.status_code != 200 or r.json().get("code") != 0:
        raise LoggingError(f"Splunk HEC rejected the event: {r.status_code} {r.text[:200]}")

The agent and the MCP server each have their own logger and share one session_id. If Splunk does not acknowledge an event within 5 seconds, LoggingError propagates and the agent exits with Agent stopped: Splunk HEC unreachable. I tested it with Splunk down: the agent never reached its first model call.

Two event families go to the HTTP Event Collector:

Family Sourcetypes Content Use
Summary, always on agent:session, agent:llm, agent:toolcall Model, tools requested, arguments, token counts, durations; question, answer and tool results stored as SHA-256 hashes Production-style audit logging without copying sensitive content
Trace, LAB_TRACE=full agent:trace One event per step in clear text: the exact context sent to the model, its reply, each sandbox decision Lab only, to read and replay a session

The summary field of each trace event is a fixed string written in the code. If the model wrote it, an injected prompt could make a malicious step look harmless in the SIEM.

What happens in one question, frame by frame?

This search rebuilds a session in order:

index=agents sourcetype=agent:trace session_id=<SESSION_ID>
| sort 0 _time
| table _time component kind step summary

Splunk search rebuilding one agent session from 25 trace events: agent start, MCP handshake, tool list, two model calls, list_files, a sandbox path check allowing incident-report-2026-09.md, then read_file

The question was “What was the root cause of the September incident?”, sent to openai/gpt-5.6-luna through OpenRouter. Here are the 25 trace events, timed from the first one. The 8 summary events of the same session (agent:session, agent:llm, agent:toolcall) are interleaved in the index and left out of this table.

# Time Component Event What happens
1 0.000 s agent agent.start The question, the model name, the system prompt and the turn limit are recorded
2 0.872 s mcp_server server.start The child process is up and logs its sandbox folder
3 0.937 s agent mcp.initialize Handshake done: protocol version and server capabilities
4 0.944 s agent mcp.tools_list 2 tools discovered, each description hashed
5 0.950 s agent llm.request Turn 1: 2 messages and 2 tool definitions sent to the model
6 3.019 s agent llm.response The model answers in 2,064 ms with one tool call, after 13 reasoning tokens
7 3.023 s agent agent.decision The loop decides to run 1 tool call
8 3.026 s agent mcp.call_tool.request tools/call list_files
9 3.031 s mcp_server server.list_files 2 files found in the sandbox
10 3.040 s agent mcp.call_tool.result 43 characters back: the two file names
11 3.046 s agent llm.request Turn 2: 4 messages sent
12 4.243 s agent llm.response The model answers in 1,191 ms with two tool calls
13 4.247 s agent agent.decision The loop decides to run 2 tool calls
14 4.249 s agent mcp.call_tool.request tools/call read_file on incident-report-2026-09.md
15 4.253 s mcp_server server.path_check Path resolved inside the sandbox: allowed
16 4.260 s mcp_server server.read_file 625 of 625 bytes read
17 4.266 s agent mcp.call_tool.result 623 characters back: one dash in the file takes three bytes in UTF-8
18 4.278 s agent mcp.call_tool.request Second tools/call read_file, on access-policy.md
19 4.282 s mcp_server server.path_check Path resolved inside the sandbox: allowed
20 4.286 s mcp_server server.read_file 360 of 360 bytes read
21 4.295 s agent mcp.call_tool.result 360 characters back
22 4.299 s agent llm.request Turn 3: 7 messages sent
23 5.697 s agent llm.response The model answers in 1,395 ms with text and no tool call, finish reason stop
24 5.703 s agent agent.decision No tool requested: the answer is final
25 5.705 s agent agent.final Session completed after 3 turns

The final answer, 217 characters: “The root cause was that the invoice assistant’s mail-sending tool had no recipient allow-list. A supplier email exploited this by including hidden instructions to forward invoices to a new external accounting address.” The incident report is fictional, written for the lab.

What the timings show:

  • The model is most of the wait. Its three answers took 4,650 ms of the 5.7 seconds the session lasted.
  • Starting the MCP server is the slowest local step: 872 ms, the time to launch a second Python interpreter. After that, each tool call made the round trip in 12 to 15 ms.
  • The model asked for two files at once at turn 2, the access policy included, although the question was only about the incident. The loop ran both calls one after the other, each with its own sandbox check. What the agent reads is bounded by its tools, so the sandbox has to be as narrow as the task.
  • The model never touched a file. Every read went through tools/call, server.path_check and server.read_file, in that order, with a log line at each stage.

Which attacks does each frame expose?

Read the frames as an attacker would, and each one points to a way of pushing the agent off course, plus the event that would show it:

Frame What an attacker targets Risk (OWASP Top 10 for LLM 2025) Signal in the logs
mcp.tools_list The tool descriptions, rewritten by a compromised or malicious server Tool poisoning, LLM01 and LLM03 description_sha256 differs from the baseline
llm.request A document returned at an earlier turn that carries hidden instructions Indirect prompt injection, LLM01 A tool sequence that differs from the usual pattern
mcp.call_tool.request The path argument: ../.env, /etc/passwd, a symlink Excessive agency, LLM06 server.path_check with inside_sandbox=false
llm.request count A loop that never reaches a final answer Unbounded consumption, LLM10 status=max_steps_reached
Summary events Secrets copied into the SIEM Sensitive information disclosure, LLM02 Hashes instead of content

I ran the path attacks against this server: ../.env, ../../../etc/passwd, /etc/passwd and a symlink planted in the sandbox. All four were refused. _safe_path resolves the path first and checks it second, which catches two cases a string check misses. In Python’s pathlib, joining a base with an absolute path returns the absolute path, and resolve() follows symlinks to where they point.

How is the lab isolated from the rest of the network?

The agent runs on its own VM, on a Proxmox bridge that reaches the internet through NAT and drops every private range. The first version of my rules had a gap that only showed up when I tested from inside the agent VM: the hypervisor’s web UI (8006) and SSH (22) answered.

Two causes:

  • Traffic to the host itself goes through the INPUT chain, so a FORWARD rule never sees it.
  • My DROP rules only matched traffic leaving through the uplink (-o vmbr0), so traffic to another bridge passed.

The fixed rules, now in the bootstrap script:

post-up iptables -I FORWARD -s 10.10.10.0/24 -d 10.0.0.0/8 -j DROP
post-up iptables -I FORWARD -s 10.10.10.0/24 -d 172.16.0.0/12 -j DROP
post-up iptables -I FORWARD -s 10.10.10.0/24 -d 192.168.0.0/16 -j DROP
post-up iptables -I FORWARD -s 10.10.10.0/24 -d 10.10.10.0/24 -j ACCEPT
post-up iptables -I INPUT -i vmbr1 -j DROP
post-up iptables -I INPUT -i vmbr1 -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT

-I inserts at the top of the chain, so the last line written is evaluated first: the intra-lab ACCEPT wins over the 10.0.0.0/8 DROP, and replies win over the INPUT DROP. Retested from inside the agent VM with ping and TCP connects, the model API and Splunk answer, and the hypervisor, the LAN and the other lab bridges do not. The Proxmox network documentation describes the masquerading setup this builds on.

How do you rebuild the lab yourself?

You need a Proxmox VE host with 9 GB of free RAM for the default sizing (4 GB for the agent VM, 5 GB for Splunk), an SSH key pair, an OpenRouter API key and the Splunk Enterprise .deb from your Splunk account. Then three steps:

  1. On the Proxmox host, as root: SSH_PUBKEY_FILE=/root/lab_key.pub bash proxmox/bootstrap-lab.sh. It creates the isolated bridge, verifies the SHA-512 of the Debian 13 cloud image and starts both VMs with cloud-init.
  2. On the logs VM: sudo SPLUNK_DEB=/tmp/splunk.deb AGENT_IP=10.10.10.10 bash splunk/configure-splunk.sh. It installs Splunk, creates the agents index, enables HEC and prints the token.
  3. On the agent VM: copy agent/env.example to ~/lab/.env, fill in the keys, run bash agent/setup.sh, then python agent.py "your question", and paste the printed session_id into the search above.

The sample documents are fictional. Trace mode writes prompts and tool results to Splunk in clear text, so keep real data out of this lab. splunk/searches.md holds six starter searches, including the tool-description baseline.

What does this lab not cover yet?

One agent, two read-only tools, a local stdio MCP server. It does not yet test an agent with write or send permissions, a remote MCP server with authentication, agent identities in an enterprise directory, or a prompt injection hidden in a document the agent reads. Those are the next episodes of the CyberAI Security Journey, each added to the same repository with its controls and detections.

Sources

Questions

How do AI agents work?

An AI agent is a loop around a language model. The loop sends the model the conversation and a list of tools. The model returns an answer or a JSON request to call a tool; the loop runs the tool, adds the result to the conversation and asks again, until the model gives a final answer or a turn limit is reached.

How does MCP work under the hood?

The Model Context Protocol is JSON-RPC 2.0 between an AI application and a tool server. With the stdio transport, the application starts the server as a child process and exchanges one JSON object per line over its standard input and output: `initialize`, then `tools/list` to discover the tools, then `tools/call` to run one.

Do AI agents remember previous steps?

The model does not. The loop keeps the history and resends all of it at every turn, which is why the context grew from 2 to 4 to 7 messages in the session above, and why token costs grow with every tool call.

Does the LLM execute the tools itself?

No. The model returns the tool name and its arguments as JSON. The agent's code decides whether to run it and sends the call to the MCP server, which executes the function. Every security control in this lab sits in that code, not in the model.

Can you build an AI agent from scratch in Python without a framework?

Yes. This one uses the MCP Python SDK for the tool server, an OpenAI-compatible client for the model, and a `for` loop. The whole agent is 278 lines across three files, logging included.

Is it safe to put real company documents in the lab?

Not with trace mode on. `LAB_TRACE=full` stores the full context and tool results in clear text in Splunk. Use the fictional sample documents, or turn tracing off and keep only the hashed summary events.