Dekita

Your agent did something — but nobody can say what. A practical audit trail pattern for agent tool calls

ai agents observability engineering

Your agent just finished a task. It called five tools, retried one, asked a human to approve another, and produced a file you didn't ask for. You want to know: which tool was called when, with what arguments, and what did it return?

If your agent runtime is anything other than a toy, the honest answer is "I don't know, and it's not easy to find out."

That's the gap this post addresses. I'll walk through a minimal, copy-paste-able pattern in Node.js for a structured audit trail that records every tool call, approval, retry, and final outcome. No framework, no SDK — just an event emitter and a file sink you can grep later. By the end you'll have a pattern that took about an hour to implement and has already paid for itself in one post-mortem.

The problem in one sentence

Most agent runtimes today log user-facing state: "Agent is thinking," "Calling tool X," "Done." They don't log the decision chain: why the model chose that tool, what arguments it passed, whether an approval gate blocked it, what the tool returned, and whether a retry happened. When something goes wrong at 3 a.m., you're left reading a wall of "Agent is thinking" lines with no way to answer "which of the five tool calls actually caused the bad state?"

What an audit event actually needs

Five fields cover 90% of post-mortems:

{
  ts:        ISO-8601 timestamp
  run_id:    one id per top-level agent invocation
  event:     "tool_call" | "approval" | "retry" | "outcome"
  tool:      tool name, or "human" for approval events
  payload:   { args, result, error, duration_ms }
}

That's it. run_id is the one thing people usually forget, and it's the reason "which of my three concurrent agents did this?" becomes unanswerable. The whole pattern below hinges on that one field.

The pattern in 30 lines

const fs = require("node:fs");
const path = require("node:path");

function makeAuditSink(file) {
  fs.mkdirSync(path.dirname(file), { recursive: true });
  return (runId, event, tool, payload) => {
    const line = JSON.stringify({
      ts: new Date().toISOString(),
      run_id: runId,
      event,
      tool,
      payload,
    });
    fs.appendFileSync(file, line + "\n");
  };
}

const audit = makeAuditSink("./audit.log");

// In your agent loop, before calling a tool:
audit(runId, "tool_call", tool.name, { args: toolArgs });
const t0 = Date.now();
const result = await tool.run(toolArgs);
audit(runId, "outcome", tool.name, {
  result: result.slice(0, 500),
  duration_ms: Date.now() - t0,
});

For retries and approvals, emit the same event type with a distinguishing field:

audit(runId, "retry", tool.name, { reason: "timeout", attempt: 2 });
audit(runId, "approval", "human", { blocked: true, policy: "no-write-prod" });

That's the whole pattern. No abstraction layer, no framework. You can bolt it onto any runtime — a hand-rolled loop, LangChain, OpenAI Assistants, whatever — because it only knows "tool name" and "payload," not what your tool actually is.

Wiring it into a realistic agent loop

Here's a more complete version that shows where the events actually land in a retry loop:

async function runAgentWithAudit(runId, tool, initialArgs, { maxRetries = 2 } = {}) {
  const t0 = Date.now();
  audit(runId, "tool_call", tool.name, { args: initialArgs, attempt: 1 });

  for (let attempt = 1; attempt <= maxRetries + 1; attempt++) {
    try {
      const result = await tool.run(initialArgs);
      audit(runId, "outcome", tool.name, {
        result: result.slice(0, 500),
        duration_ms: Date.now() - t0,
        attempt,
      });
      return result;
    } catch (err) {
      if (attempt > maxRetries) throw err;
      audit(runId, "retry", tool.name, {
        reason: err.message.slice(0, 200),
        attempt,
      });
      await new Promise(r => setTimeout(r, 1000 * attempt));
    }
  }
}

Notice the attempt field on the outcome. Without it, a retried call and its final success look identical in the log, and you lose the signal "this tool was flaky in this run." That signal is what you actually want when you're deciding whether to raise the retry cap or fix the underlying tool.

Reading it back

Once you have the log, post-mortems become a grep:

# All failed tool calls in one run
grep '"run_id":"run-abc-123"' audit.log | grep -E '"event":"retry"|"event":"outcome"' \
  | jq -r 'select(.event=="outcome") | .payload.result'

# Which tools ran longest this week?
jq -r 'select(.event=="outcome") | [.tool, .payload.duration_ms] | @tsv' audit.log \
  | sort -k2,2n | tail -10

# How often did the "no-write-prod" approval gate actually block something?
grep '"policy":"no-write-prod"' audit.log | wc -l

The third one is the surprising one. I expected the gate to block almost nothing, because we'd tuned it to be permissive. It blocked ~1 in 40 runs, but every single block was the agent trying to rm -rf a path it had hallucinated. The log told us that in one line; without it, that would have been a silent production risk.

Two things I'd not do

  1. Don't log secrets in args. If a tool takes a token, redact it in the payload before writing. The audit log outlives the process — don't make it a second copy of your credentials. A one-line helper is enough:
   function redact(obj, keys = ["token", "apiKey", "secret"]) {
     const out = structuredClone(obj);
     for (const k of keys) if (k in out) out[k] = "***";
     return out;
  }
  1. Don't make the audit format depend on the agent framework. The schema above is deliberately boring: five fields, JSON lines. That means you can switch runtimes without re-training anyone to read the log. I've seen teams build rich framework-specific audit formats that became orphaned the day they changed frameworks. The boring format is the one that survives.

Where this lands on the four-gates framing

If you've read the "Four Gates" framing — access, action, approval, audit — the audit event I'm describing is Gate 4, the one that lets you reconstruct what happened after the fact. The other three gates are hard to retrofit once the agent is in production; audit is cheap to retrofit because it's just writing lines to a file. That makes it the highest-ROI gate to add first. Start there, then work backwards.

The approval gate is where the log pays for itself

The approval event type is the one most people skip, and the one that's hardest to add later. When your agent has write access to anything, you'll eventually want a gate that says "this action needs a human." The audit line for that gate is cheap:

audit(runId, "approval", "human", {
  tool: "shell",
  args: { cmd: "rm -rf ./dist" },
  blocked: true,
  policy: "no-write-prod",
  decided_by: "jane@corp",
});

The decided_by field matters more than you'd think. Six months from now, when a regulator or a post-mortem asks "who authorized this write?" you want a name in the log, not a guess. The gate itself is usually two lines in your runtime (check a policy, block if it fails); the record of the check is what the audit gives you. Without it, the gate is a black box — you can tell it blocked something, but not who decided the policy or what the rejected args were.

A note on scale

JSON-lines on a local file works fine until you're generating more than a few thousand events per minute. Past that, you'll want to push the same event shape to a real sink (S3, CloudWatch, a log store) — but keep the shape identical. The whole value of the pattern is that the schema outlives the transport.


This post was written with AI assistance. The author is responsible for its content.