For teams that care about agent performance

Your coding agents at full speed in your codebase

Scott simulates thousands of coding-agent sessions in your repo to find and fix what slows them down. Early teams see 15%+ agent speedup in under two weeks.

scott drag — baseline

$scott drag --baseline

>simulating 10,000 sessions

>10/31 findings resolved

>8.9% faster agents, 10.4% less token usage

The Scott Report in Action

One button press -> ranked report -> pull request that resolves the finding

Built by Coinbase engineers, backed by YC, MaC Ventures and world class operatores

Y Combinator

Fall 2025

Ex-Coinbase

Dev platform - 700+ engineers

MaC Ventures

Venture fund

Garry Tan

President & CEO, YC

Agent drag is sneaky expensive

Your repo, not the model, is the biggest variable in how well agents perform. Most of that cost is invisible until you measure it.

01

Agents thrash on your codebase

The same agent that flies on a clean repo stalls on yours: retrying, re-reading, second-guessing. That drag is invisible on a dashboard but shows up in slow PRs and runaway token bills.

02

You can't see how agents walk your code

Public benchmarks don't predict performance on your repo. What actually slows agents down — flaky tests, conflicting docs, tangled structure — is specific to your codebase and impossible to read from the outside.

03

Every model release resets the game

Agent behavior on your repo shifts silently with each model upgrade. What worked in June behaves differently in July, with no way to catch the regression before it costs you.

How it works

Simulate, diagnose, resolve. Your team stays in the loop on every change.

Simulate

Thousands of coding-agent sessions run against a sealed, read-only copy of your repo: realistic tasks, pinned models, full trajectories captured.

Diagnose

Every session is measured and ranked. Scott surfaces the specific drag slowing your agents down, ordered by how much speed and token cost each fix is worth.

Resolve

For the high-value, low-risk items, Scott opens a ready-to-review pull request. The rest ship as a ranked fix plan. Your engineers review, approve, and merge — always in control.

What the baseline finds

A ranked report of what's slowing your agents down, and where your repo stands against anonymized peer codebases.

Flaky test suites

agents retry and stall on nondeterministic tests

Missing context anchors

undocumented invariants force repeated re-reading

Tangled module boundaries

wide blast radius makes agents cautious and slow

Broken local dev setup

agents waste turns getting the repo to run

Sample drag items — your report is generated from your repo.

Why the numbers hold up

Research-grounded

Built on leading-edge research into agent harness performance and evaluation.

Hard to replicate

A simulation platform that measures improvement on your repo, not a generic benchmark.

Model-agnostic

Runs across open and private models, so improvements hold on Claude Code, Cursor, Devin, Copilot, and whatever ships next.

Next in the Scott suite

Once your repo is in order, uplevel your team with Scott Compass.

Spend less time on code review and focus on design review so your agents build the right thing →

Get started

See your repo's drag.

We'll run the baseline on your codebase and walk you through exactly what's slowing your agents down.