How Google tests billions of lines of code — without running every test
Google's monorepo contains over 2 billion lines of code. Every day, engineers make tens of thousands of changes to it. And every one of those changes needs to be tested before it can land.
Running the full test suite for every change isn't just slow — it's physically impossible. So Google built something smarter. They call it TAP, and at its core is an idea that every engineering team can apply right now: Test Impact Analysis.
The problem with running everything
Most teams start with a simple rule: any change triggers the full test suite. It feels safe. It feels thorough. And for a while, it works fine.
Then the codebase grows. The test suite grows. CI starts taking 20 minutes, then 40, then over an hour. Engineers start ignoring red builds because there are too many. The feedback loop — the thing that makes testing valuable — slowly breaks down.
Google hit this wall earlier than almost anyone. At their scale, running every test for every change would require an amount of compute that makes the problem trivially unsolvable. Their engineering productivity team had to find a fundamentally different approach.
TAP: Google's Test Automation Platform
TAP (Test Automation Platform) is Google's internal CI system. It processes hundreds of millions of test executions per day and acts as the gatekeeper between a code change and the main branch. Nothing lands without passing TAP.
But TAP doesn't run everything. It runs the right things.
The system works in two modes. Presubmit testing runs before a change is committed — catching regressions before they ever reach the shared codebase. Post-submit testing continuously validates the health of the mainline. Both rely on the same underlying mechanism: knowing which tests matter for a given change.
How Google's Test Impact Analysis works
Google's TIA is built on top of their build system — Blaze internally, open-sourced as Bazel. Every target in the build graph has explicit, declared dependencies. When a file changes, the system walks the dependency graph to find every test that depends on that file, transitively.
If auth/token.py changes, TAP doesn't guess which tests might be affected. It knows — because the dependency graph tells it exactly which test targets include auth/token.py in their transitive closure.
Change arrives
An engineer submits a changelist (CL). TAP receives it and identifies every file that was modified.
Dependency graph traversal
Blaze/Bazel's dependency graph is queried. Every build target that transitively depends on the changed files is identified.
Affected test selection
From the affected targets, only test targets are extracted. These are the tests that could be broken by this change — and only these are scheduled to run.
Prioritisation
Tests are ordered by historical failure rate, recency of the files involved, and number of recent authors — to surface real regressions as quickly as possible.
Results + feedback
Results are returned to the engineer. Fast tests first. The goal is actionable feedback in minutes, not hours.
Static analysis vs. runtime coverage
Google's approach — using declared build graph dependencies — is a form of static analysis. It's extremely fast because there's no need to instrument test runs: the dependency information already exists in the build system.
This works well when dependencies are explicit and complete. In a strictly-typed, monorepo-native build system like Blaze, that's mostly true. Every import and dependency must be declared in a BUILD file.
Python, however, is a different world. Dependencies are often implicit. Dynamic imports, plugin systems, monkey-patching, and subprocess calls mean the static dependency graph frequently under-counts what a test actually exercises at runtime.
That's why runtime coverage-based TIA — which is what DeltaTest uses — tends to give more accurate results in Python codebases. Instead of inferring what a test might touch, it measures what a test actually executed, line by line, including subprocesses.
What else Google does that most teams skip
Hermetic tests
Google invests heavily in hermeticity — tests that are self-contained, isolated, and deterministic. A hermetic test doesn't talk to a real database, doesn't depend on environment variables, doesn't share state with other tests. This is what makes it safe to run a subset: you know a test's result depends only on the code it's testing, not on global state left by a previous test.
Flaky tests undermine TIA. If a test sometimes fails for reasons unrelated to the code change, the system can't trust its signals. Google has entire infrastructure dedicated to detecting and quarantining flaky tests automatically.
Test ownership and size classification
At Google, every test has a declared size: small, medium, or large. Small tests run in milliseconds and have no external dependencies. Large tests can take minutes and may touch real services. TIA prioritises small tests in presubmit — they're fast, reliable, and cheap. Larger tests run post-submit or on a schedule.
This classification is enforced at the build system level, not just a convention. It's why Google can confidently run a focused subset and trust the results.
Milestone batching
For post-submit testing, TAP doesn't run after every single commit. It batches consecutive commits into milestones — snapshots of the codebase taken approximately every 45 minutes during peak hours. This amortises the cost of large test suites across multiple changes, balancing cost against the risk of slightly delayed feedback on mainline health.
The lesson for every engineering team
You don't need a monorepo, a custom build system, or 1,000 SREs to apply the core idea. The principle is the same whether your codebase has 500 files or 50 million:
- Don't run tests that can't possibly be broken by a change. Every unnecessary test run is wasted compute and wasted time.
- Build a dependency map between your code and your tests. Whether static or coverage-based, this map is what makes test selection possible.
- Prioritise fast feedback over exhaustive coverage. A test suite that runs in 2 minutes and catches 95% of issues is more useful than one that runs in 90 minutes and catches 100%.
- Invest in test quality. TIA amplifies good tests and amplifies bad ones too. Flaky, non-hermetic tests will produce unreliable signals at any scale.
pip install pytest-deltatest.
TIA for Python teams, today
DeltaTest brings the same Test Impact Analysis principle to pytest — without requiring a monorepo, a custom build system, or a dedicated infrastructure team.
Instead of declaring dependencies in BUILD files, DeltaTest instruments your existing pytest runs with a lightweight coverage plugin. It builds a precise, line-level map of which tests exercise which lines of code — including subprocesses. That map is synced to the cloud so your whole team benefits from each other's runs.
On the next run, only the tests whose mapped lines overlap with your git diff actually execute. Everything else is skipped — not because we guessed it was safe, but because we measured it.
The result is a feedback loop that's fast enough to run on every commit, every pre-commit hook, every PR — not just in CI once a day.
Run the 3% that matters
Install in 60 seconds. Works with your existing pytest setup. No build system changes required.
Get started free →Sources: Software Engineering at Google (Winters, Manshreck, Wright, 2020) · Google Testing Blog · bazel.build