A contact sheet: sixteen timestamped frames from a rendered test clip, showing a character climbing a flight of concrete stairs in a grey test level.

How a coding agent verifies a 3D game it can't watch

How I get an AI coding agent to check a 3D game it can't watch, and the day that process passed a broken build three times.

A large share of the code on the game I'm working on is written by an AI coding agent, Claude Code. That raises an obvious problem: the agent can't watch the game. It can look at images and it can read logs, and that's it. So the agent and I built a process that lets it check its own work using only those two things. This post walks through how that process works, and then through the day it told me a broken build was fine three separate times.

It helps to know what I actually want out of it. The point isn't to make the game correct. It's so I know what state the build is in and can decide what to fix. Most of what it finds, I leave alone for now. The game is several months away from having players, and at this stage a wrong call is cheap to undo. A problem I know about and chose to defer means the process did its job. What I'm trying to avoid is a problem I don't know about at all, because then I never got to make that call. Every incident below is one of those.

Nothing here reached a player. The broken build sat on an integration branch for a day before I found it.

The game's repository is private, but the parts worth looking at are in a public Claude Code plugin: a skill for writing and rendering test drivers (godot-movie-driver), a skill for turning a video into images an agent can read (video-review), and the definition of a second, blind reviewer agent (frame-review). There's no generative art on the project, and the art team's work isn't in the build yet.

Why numbers and screenshots aren't enough

Three things make a 3D game hard for an agent to check, and none of them are unique to this project.

The first is that bone data doesn't tell you what was drawn. In Godot, Skeleton3D.get_bone_global_pose() returns the pose from before any SkeletonModifier3D runs, and an IK solver is a SkeletonModifier3D. So while an IK solver is visibly bending a leg, that call is still handing you the rest pose. Two of the earliest bugs on the project showed us this: a character that rendered buried up to its waist, and a walk where the feet never touched the ground. No numeric check caught either of them. One rendered frame showed both.

The second is that a screenshot is a single moment, and most of the questions I care about are about something changing. Does the animation blend finish, does the foot stay planted, does the weapon jump somewhere it shouldn't? If the question has a verb in it, a still image can't answer it.

The third is the one this post is really about. The game is built to hide its own failures. If the animation clips fail to load, it falls back to a procedural walk. A missing rifle animation set leaves the character unarmed, and a missing death animation leaves the body frozen in place. Those are all good engineering decisions, but every one of them still leaves you with a character that walks around, so an agent asking "does the character walk and animate?" will say yes.

How the process works

A driver plays the game

A driver is a small GDScript node the agent writes for one test and usually throws away afterwards. It's a state machine running in _physics_process. It teleports to the scene it's testing, presses and releases inputs on timers, reads game state between phases, prints what it saw, and quits. The only way to turn one on is a command-line flag:

godot --path game -- --zone=gym --driver=WeaponSocketSweep

The main scene reads --driver=<name> and, in a debug build, loads res://TestDrivers/<name>.gd as a child. Release builds ignore the flag, and if the driver doesn't exist the game logs a warning and carries on.

We tried a different approach first, with the agent adding a temporary autoload to the project file, and it caused two bugs on the first day. An open Godot editor rewrites the project settings from its in-memory copy, so entries the agent had removed from disk kept coming back. The flag doesn't touch any file, so that can't happen.

Movie Maker renders every frame

The driver runs under Godot's built-in Movie Maker mode, which renders every frame offline at a fixed timestep:

dotnet build game/Game.sln    # the render loads the on-disk assembly, not the editor's
godot --path game --write-movie /abs/out.avi --fixed-fps 60 -- --zone=gym --driver=foo

The same driver on the same build gives you the same video every time, and nobody has to be at the keyboard. Two things are worth knowing before you copy this. It always records at the project's design resolution, and nothing you do at runtime changes that. It also won't run with --headless, because that flag picks a renderer with no framebuffer and the process crashes on the first frame it tries to record. On a Linux desktop with the screen locked, setting DISPLAY=:0 is enough.

The agent reads the log before the frames

When you watch a video you see all of it. An agent sees a handful of sampled frames, and anything that happens between two samples is invisible to it. The log doesn't have that gap. The driver writes to it at the physics rate, and it records what the state machine did, which isn't always what ended up on screen. So every driver prints its state changes with timestamps and the values that caused them:

t=2.53  tapped move_left (2 frames)
t=2.55  idle_r -> walk   (speed=1.50)
t=2.57  walk   -> run    (speed=3.00)
t=2.60  run    -> idle_l (speed=0.00)

That excerpt is from the day this became a rule. The agent was checking whether a 0.2-second crossfade between two idle poses looked wrong. The frames showed a turn that looked fine, but the log showed the character's speed crossing the idle threshold within a single physics frame. The crossfade it was testing couldn't be triggered from the keyboard at all. Going by the frames alone, it would have concluded the crossfade looked fine, when in fact it had never run.

Contact sheet first, then the full frame

Since the agent can't play a video, it turns the video into images in two steps. First comes a contact sheet: about sixteen timestamped frames tiled into one image, kept at or under the width the model can see without it being scaled down. For the current model that's roughly 1500 pixels. A 4K sheet is actually worse, because it gets shrunk before the model sees it and the small text turns to mush. The sheet is for finding when something happened. It's too small to show you what happened, so the second step is pulling a full-resolution frame at that timestamp and cropping it to the part that matters.

It's almost always worth cropping to the subject before tiling. Sixteen uncropped 1080p frames in one sheet leave the character a few pixels tall.

A second reviewer that only sees pictures

When the agent running the test shouldn't trust its own reading of the frames, it hands them to a second agent with a fixed contract. The dispatching agent has to give it four things, and it refuses to review without them: the video or the extracted images, the choreography with timestamps, a description of what correct looks like, and which fallback could be hiding the problem. The reviewer doesn't read logs, code, or notes. It lists its findings, most serious first, with the frames it's basing each one on, and then commits to one of two verdicts: the build looks visually correct for the choreography, or something is wrong and here's what. It's specifically not allowed to turn a persistent anomaly it can't explain into a request for a better render. I'll get to why further down.

I make the final call

What comes out of all this is evidence: a video I can watch later, a log timeline, findings tied to specific frames, and a verdict from a reviewer that only saw the pictures. None of it is a sign-off. An agent finishes its work with "Changes made, ready for testing when you are." I'm the one who says "Verified as fixed."

The day it passed a broken build three times

A pull request re-exported all 51 animation clips with a setting that moved each clip's first keyframe from 1/30 of a second to zero, which made every clip one frame shorter. It re-measured the nine speed constants that depend on clip timing, but it missed one other thing that depended on it: the hard-coded times that split the jump clip into its rise, apex, and fall. One of those times was now 0.867 seconds, in a clip that was only 0.833 seconds long.

The animation loader has a guard for bad clip data, and it did exactly what it was built to do. It logged one warning, switched off all clip-based animation (unarmed, rifle, and death sets together), and dropped the game back to the fully procedural walk. The character kept walking, just more stiffly.

Here's how that got past three rounds of checking.

  1. The pull request that caused it merged with the flag "STILL NOT VERIFIED BY RENDER," because the machine only allows one game run at a time and another session had it. The flag was accurate, but the render it put off never happened.
  2. The next pull request rendered its own changes on a branch cut before the break and verified its weapon-socket geometry there to within a few microns. Then it merged onto the broken integration branch, where that evidence no longer applied.
  3. The third pull request rendered the broken build, and in its render the weapon showed up as a smudge pointing at the camera. The agent measured it. After temporarily removing its own changes, it saw the smudge was still there, so it knew its changes hadn't caused it. It filed the problem as "PRE-EXISTING, NOT INTRODUCED HERE" and merged. It was right that it hadn't caused the problem; the mistake was stopping there. The smudge was a symptom of the real failure. With the rifle clips switched off, the weapon's position was being worked out from unarmed, procedural hand poses.

On top of that, every render log that day contained the line "clip locomotion is disabled," and no session searched its log for warnings before looking at frames.

The next day I noticed the character was moving stiffly, and the smudge the third pull request had filed turned out to be the lead that explained it. I know the game, so I saw it straight away. An agent checking whether the character walks and animates won't, because the fallback was designed to be a character that walks.

We added five rules after that.

  • Search the render log for WARNING before looking at a single frame. If a warning names the system under test, the video is showing the fallback and there's nothing to review.
  • A driver has to check that the system it's testing is actually active. Checking that its measurements are in range isn't enough. One that only checks its own numbers can pass while it's filming the fallback, and the numbers will look close to perfect because it's comparing the fallback to itself. In practice that's an expect_anim check per phase: this phase has to see a clip whose name contains walk, or it fails.
  • Evidence from a branch expires when the branch merges. If other changes to the same subsystem landed while the branch was open, render a quick check again on the merged result.
  • "Not verified by render" is an unfinished task. The first session that's able to render does it before building anything on top.
  • "Pre-existing, not introduced here" narrows the scope of your change. It doesn't close the defect. Tracking down the cause with git bisect and a render is cheap, and here the "pre-existing smudge" led straight to the real failure.

The experiment that followed

Three days later I wanted to know whether a different model would have caught it. The agent re-rendered the broken build and a healthy one with the same driver, turned each render into the same kind of evidence packet (a contact sheet cropped to the character plus five cropped full-resolution frames, with no logs), and gave each packet to blind reviewers that only saw one build. We used two Claude models, so there were four reviewers in all. These were the models available in August 2026, so read the result as being about those two and not their successors.

On the broken build, the reviewer running Claude Fable 5 said "something is wrong with this build, specifically the weapon," with high confidence, and described the smudge as a degenerate mesh that never pointed at the crosshair. The reviewer running Claude Opus 5 saw the same anomalies, rated them medium confidence, suggested occlusion as a sufficient explanation, and concluded "nothing in these images indicates a broken build." That's the same pattern as the original failure: it measured the symptom and didn't draw the conclusion.

This was one sample per cell and one scenario, and both reviewers were nudged the same way by the context they were given. Treat it as an anecdote. It's still the reason the reviewer agent works the way it does: it has no access to logs or code, it's pinned to the model that gave a verdict, it has to give one, and it can't swap an anomaly it can't explain for a request for a better render.

It's held up since. A few weeks later, a reviewer caught something a fully instrumented driver had passed. An enemy sat in its chase state at zero velocity for most of a second while its walking animation kept playing, so it was walking in place. Every state check in the driver passed, because the problem was a mismatch between the animation and the velocity, and a state machine doesn't look at that. The reviewer measured the gap between the enemy and a nearby prop across several full-resolution frames and showed it didn't change while the legs moved. There's now a cheap driver check for "stationary in a state whose animation is a locomotion animation," but we only knew to write it because of the pictures.

The same mistake, producing a false failure

The incident above gave us a pass we shouldn't have trusted. The same kind of mistake can give you a failure you shouldn't trust, and that's no safer.

One of the agent's drivers measured stride, in two passes. The first froze an enemy and played its walking animation directly at rate 1.0, to measure how fast the animation moves across the ground. The second turned the enemy's state machine back on and read the playback rate the game chose. That second pass was the one that mattered, and it couldn't work, because the driver never let go of the animation player. The game only starts an animation if it isn't already playing, and the walking animation was already playing at the rate the driver had set, so the game never restarted it.

So the second pass read a rate of 1.000, reported "SYSTEMATIC mismatch 0.423 m/s (22.4%)," and failed every single run for as long as the driver had existed. The rate the game actually uses is 0.816. Once the agent changed the driver to release the animation player, it reported 0.001 m/s and passed. The game had never had a problem.

The rule we took from it: if a driver controls a subsystem directly and then wants to see how the game uses it, it has to explicitly let go first. A good way to check is to ask what you'd read if the game never touched it again. If the answer is "the value the driver set," the driver is measuring its own input. A wrong failure costs you sessions chasing a bug that isn't there, and it teaches you to ignore the driver, including later when it reports something real. This one had already been written into the issue tracker as a known property of the build.

A check that can't fail

This one is the hardest to spot, because the output looks like unusually good evidence.

A headshot driver needed to sample points in a "head band," and it did that by aiming a ray at the head bone. The head bone is the center of the sphere the game uses to decide whether a hit is a headshot, so the ray went through the center of that sphere every time, by construction. The measured distance from the center was always exactly zero, the band printed head-aim 0.000 .. 0.000 -> 100% (want 100%), and a sweep across different sphere sizes scored 100% at every size, including one centimeter. The agent reported that sweep as evidence ("no trade curve left to pick a row off"), but the check couldn't have come out any other way.

For every threshold a driver checks, it's worth asking what value would make the check fail. If the answer is "nothing would," the sample is broken. A sweep that scores the same at every row should make you suspicious, because it usually means the check can't fail. The fix that works in general is to build the probe from something measured independently of the value you're testing. Here that was half the height of the head's bounding box, taken from the mesh, and after that change the sweep started failing at small sizes the way it should.

Where this came from

None of this started as testing. It started with trying to get the agent to learn inverse kinematics and rigging from tutorial videos. The parts of a technique that nobody explains out loud are only visible in the video. The transcript records what the presenter says they're doing, while the frames show what they actually do, and the two often differ. In one video the presenter says a feature is "included in Blender" while the screen shows them installing a third-party fork, and automatic captions regularly get keyboard shortcuts wrong. So the study process read the transcript first to decide where to look, and then went through every segment of the video. It's the same idea as reading the log before the frames: the cheap source tells you where to look in the expensive one, and neither one replaces the other.

What it costs, and what it doesn't do

A render runs somewhere between a third of real time and twice real time, so a 13-second clip takes a minute or two. The review is a handful of images. A typical check is a few minutes of agent time and a few megabytes of evidence I can look at later, and a blind second review roughly doubles that. The expensive part is writing a driver that stages the test honestly, and everything above is what it cost us to learn what honest means.

It doesn't tell me whether the game is any good. How it feels, how it's framed, and whether a walk looks heavy are judgments I make from the same video. It can only verify what a driver can set up, and when a system can't be reached from a driver, the fix is usually to make the game code more observable. It won't catch a fallback on its own either. The warning search and the engagement check are what protect against that, and they're habits, not tooling. The reviewer comparison is one controlled sample. And it needs a real display.

If you want to try this

Roughly in order of how much each one would change a typical setup:

  1. Turn test drivers on with a command-line flag instead of editing a tracked file. It's about a dozen lines in the main scene's ready method, in debug builds only.
  2. If the question has a verb in it, render a video instead of taking a screenshot. Movie Maker is built into Godot, and all it needs is a display and an absolute output path.
  3. Have the driver log with its own timestamps, and log each state change with the values that caused it. Read the log before the frames, and search it for warnings before you open anything else.
  4. Review frames in two steps: use a contact sheet to find the moment and a full-resolution frame to look at it. Crop before you tile.
  5. Have every driver check that the system under test is active, and for every threshold, ask what value would make the check fail.
  6. When the stakes justify it, use a blind second reviewer with a fixed contract and a verdict it has to commit to.
  7. Keep the final decision with a person. The process gives you information for that decision. It doesn't make it.

Items 2 through 4 are the two skills in the plugin repo, and the loader for item 1 is reproduced in the godot-movie-driver skill. The contract for item 6 is a single Markdown file, agents/frame-review.md, and if you only read one file from the repo, make it that one.