For teams that care about agent performance
Your coding agents at full speed in your codebase
Scott simulates thousands of coding-agent sessions in your repo to find and fix what slows them down. Early teams see 15%+ agent speedup in under two weeks.
Built by the engineers behind Coinbase's data layer and design system.
$scott drag --baseline
>simulating 10,000 sessions
>10/31 findings resolved
>8.9% faster agents, 10.4% less token usage
The Scott Report in Action
One button press -> ranked report -> pull request that resolves the finding
Y Combinator
Fall 2025
Ex-Coinbase
Dev platform - 700+ engineers
MaC Ventures
Venture fund
Garry Tan
President & CEO, YC
Agent drag is sneaky expensive
Your repo, not the model, is the biggest variable in how well agents perform. Most of that cost is invisible until you measure it.
01
Agents thrash on your codebase
The same agent that flies on a clean repo stalls on yours: retrying, re-reading, second-guessing. That drag is invisible on a dashboard but shows up in slow PRs and runaway token bills.
02
You can't see how agents walk your code
Public benchmarks don't predict performance on your repo. What actually slows agents down — flaky tests, conflicting docs, tangled structure — is specific to your codebase and impossible to read from the outside.
03
Every model release resets the game
Agent behavior on your repo shifts silently with each model upgrade. What worked in June behaves differently in July, with no way to catch the regression before it costs you.
How it works
Simulate, diagnose, resolve. Your team stays in the loop on every change.
Thousands of coding-agent sessions run against a sealed, read-only copy of your repo: realistic tasks, pinned models, full trajectories captured.
Every session is measured and ranked. Scott surfaces the specific drag slowing your agents down, ordered by how much speed and token cost each fix is worth.
For the high-value, low-risk items, Scott opens a ready-to-review pull request. The rest ship as a ranked fix plan. Your engineers review, approve, and merge — always in control.
What the baseline finds
A ranked report of what's slowing your agents down, and where your repo stands against anonymized peer codebases.
agents retry and stall on nondeterministic tests
undocumented invariants force repeated re-reading
wide blast radius makes agents cautious and slow
agents waste turns getting the repo to run
Sample drag items — your report is generated from your repo.
Why the numbers hold up
Research-grounded
Built on leading-edge research into agent harness performance and evaluation.
Hard to replicate
A simulation platform that measures improvement on your repo, not a generic benchmark.
Model-agnostic
Runs across open and private models, so improvements hold on Claude Code, Cursor, Devin, Copilot, and whatever ships next.
Next in the Scott suite
Once your repo is in order, uplevel your team with Scott Compass.
Spend less time on code review and focus on design review so your agents build the right thing →