Somewhere along the way, "we should have good test coverage" turned into "every pull request must increase the coverage percentage". It sounds rigorous, and it's one of the faster ways to rot a test suite, because people stop optimizing for covering the code they wrote and start optimizing for making a number go up. This post reads coverage as a trend instead, then wires that up in a TypeScript project with Vitest and Codecov on GitHub Actions (with the pytest-cov equivalent for Python). You'll leave with a codecov.yml that checks the code each PR adds and only reports on the total.
A target for the codebase, reached together
Pick a target for the codebase as a whole, say 80%, and reach it collectively over time. You don't get there by demanding that each PR push the percentage higher. You get there by making sure everyone covers what they build, and then reading the trend as a signal rather than a gate.
Three patterns to read
Coverage movement only tells you something when you compare it with the work that happened.
New code shipped with its tests: the ratio holds steady. You added code and its tests together, so the percentage stays about the same. This is the healthy default and what should happen most of the time. A flat number here is good news, not a failure to "improve".
Tests written for previously untested code: the ratio climbs. You went back and covered paths nobody exercised before. Also healthy. This is what backfilling looks like, and it's worth celebrating rather than treating as the only acceptable outcome.
New code with no tests, or a pile of tests with no movement: investigate. This is the pattern that's actually a signal. A drop means code shipped untested and pulled the number down. A batch of new tests that doesn't move the number deserves a look too. It often means they hit paths already covered, adding runtime and maintenance without adding protection (the question before you write any integration test is whether it covers a path nothing else covers).
Why the per-PR gate backfires
The moment coverage becomes a gate on every PR, you create incentives to game it. A three-line bug fix gets padded with a trivial test so the number ticks up. People avoid touching legacy code because doing so would "lower coverage" on their PR. Reviewers argue about a percentage instead of about whether the tests are any good. None of that makes the software more correct.
Worse, the gate punishes exactly the right behavior. Shipping new code with proportionate tests keeps the ratio flat, and a "must go up" gate reads flat as failed.
Project vs patch: two different questions
Codecov splits coverage into two checks that map neatly onto this. Per the Codecov status docs, the project status "measures overall project coverage and compares it against the base of the pull request", while the patch status "measures lines adjusted in the pull request". The patch check asks whether this PR covered what it built. The project check is the thermometer.
- Patch: the lines this PR changed must be 80% covered
- Project: shown on every PR,
informational: true, never blocks - The trend gets read at retro, against the work done
- Project:
target: auto,threshold: 0%, so any dip fails - A refactor that deletes covered code turns the build red
- People pad tests until the number moves
Wiring it up: Vitest, GitHub Actions and Codecov
Start with coverage output Codecov can read. Vitest's v8 provider is the default (install @vitest/coverage-v8), and the lcov reporter writes coverage/lcov.info.
vitest.config.tsimport { defineConfig } from 'vitest/config';export default defineConfig({test: {coverage: {provider: 'v8',reporter: ['text', 'lcov'],include: ['src'],},},});
Vitest also has coverage.thresholds, which fails the run when the totals drop below a number. That's a project-level gate living inside the test runner, so I leave it off and let Codecov do the comparing.
Then upload the report from CI with codecov/codecov-action:
.github/workflows/test.ymlname: teston: [push, pull_request]jobs:test:runs-on: ubuntu-lateststeps:- uses: actions/checkout@v7- uses: actions/setup-node@v7with:node-version: 24- run: npm ci- run: npx vitest run --coverage- uses: codecov/codecov-action@v7with:files: ./coverage/lcov.infotoken: ${{ secrets.CODECOV_TOKEN }}
On a Python service the only change is the test step: pytest --cov=src --cov-report=xml:coverage.xml, with files: ./coverage.xml on the upload.
Finally, the policy itself, in codecov.yml at the repo root:
codecov.ymlcoverage:status:project:default:target: auto # compare with the PR's base committhreshold: 1% # allow a 1% dip before it reads as a dropinformational: truepatch:default:target: 80% # the lines this PR touched
target: auto compares against the base commit, threshold lets coverage drop by that much and still pass, and informational: true makes the status pass whatever the number. So the project check still shows up on every PR, it just can't block one. The patch check is the only one that can go red, and it only looks at code the author wrote.
Read the trend, don't stand guard at it
Track project coverage over time and look at it the way you'd look at any health metric: does the trend match the work? New feature plus tests, ratio steady: fine. Backfill, ratio up: great. Number dropping, or tests landing with no effect: go look. That conversation is useful in a way that "did this PR raise the number?" never is. It also pairs well with tests that are worth counting in the first place, which for anything touching a database means isolating the data each test writes.
Coverage is a thermometer, not a turnstile: gate the code a PR adds, and read the total as a trend.
Why it matters for your team: on a payments codebase you want the coverage conversation to be about whether the refund path is tested, not about why a refactor of the ledger module dropped the total by 0.3%.