I've reviewed a lot of test suites that pass green while production is on fire. Almost always the cause is the same: the tests mock too much, stubbing out the very code they claim to test, so the green checkmark proves only that the mock returned what the mock was told to return. This post shows the one rule that fixes it, using a Stripe payment flow in TypeScript with Vitest and MSW, plus the same idea in pytest at the end.
One rule: mock the boundary
Mock the boundary, not the code.
Draw a line around the system you own. Everything inside that line runs for real in tests: your handlers, your services, your validation, your database queries, your permission checks, and your wrapper around the payment SDK. Everything that crosses the line into something you don't control gets mocked.
For a payment flow, the line is the HTTPS request to api.stripe.com. The Stripe SDK isn't your code, but it runs in your process, so it runs for real too. The only thing you fake is the network reply.
- Payment providers (Stripe and friends), at the HTTP call
- Other third-party APIs: KYC, FX rates, accounting
- Cloud storage reads and writes
- Email and notification sending
- Token validation against an external identity provider
- Your database queries
- Your resolvers, controllers or handlers
- Your service and business-logic layer
- Your wrapper around the payment SDK
- Your authorization logic
The trap: patching your own payments client
Say a checkout service with a PaymentsClient class that wraps stripe-node (an illustrative example, though the shape is common). The tempting test mocks that class:
tests/charge-order.patched.test.tsimport { vi, test, expect } from 'vitest'import { chargeOrder } from '../src/payments/charge-order'vi.mock('../src/payments/payments-client', () => ({PaymentsClient: class {charge = vi.fn().mockResolvedValue({ id: 'pi_123', status: 'succeeded' })},}))test('charges the order', async () => {const result = await chargeOrder('order-1')expect(result.status).toBe('succeeded') // true because we said so})
A test exists to catch a regression in your code. The moment you mock your own client to "return a payment intent", you've replaced the thing under test with a puppet. It can never fail, so it can never warn you.
The puppet also hides the code that matters most. The wrapper is where $20.00 becomes 2000 (Stripe takes amounts in the smallest currency unit), where the currency gets lowercased (Stripe expects a lowercase ISO code), and where a declined card turns into your own domain error. None of that runs.
The fix: answer for Stripe at the network
MSW intercepts outgoing requests in Node, so stripe-node makes a real HTTP call and MSW answers it. Watch the imports: MSW 3.0 moved http and HttpResponse to msw/http and renamed onUnhandledRequest to onUnhandledFrame (see the 2.x to 3.x migration guide). Setting it to 'error' means any call you forgot to mock fails loudly instead of reaching the real API.
tests/charge-order.test.tsimport { setupServer } from 'msw/node'import { http, HttpResponse } from 'msw/http'import { beforeAll, afterEach, afterAll, test, expect } from 'vitest'import { chargeOrder } from '../src/payments/charge-order'import { createOrder, getOrder } from './factories'const stripe = setupServer()beforeAll(() => stripe.listen({ onUnhandledFrame: 'error' }))afterEach(() => stripe.resetHandlers())afterAll(() => stripe.close())test('a declined card leaves the order unpaid and records why', async () => {let sent: URLSearchParams | undefinedstripe.use(http.post('https://api.stripe.com/v1/payment_intents', async ({ request }) => {sent = new URLSearchParams(await request.text())return HttpResponse.json({error: {type: 'card_error',code: 'card_declined',decline_code: 'insufficient_funds',message: 'Your card was declined.',},},{ status: 402 },)}),)const order = await createOrder({ total: '20.00', currency: 'USD' })await chargeOrder(order.id) // handler, service, wrapper, SDK: all realexpect(sent?.get('amount')).toBe('2000')expect(sent?.get('currency')).toBe('usd')const saved = await getOrder(order.id)expect(saved.status).toBe('payment_failed')expect(saved.failureReason).toBe('insufficient_funds')})
(createOrder, getOrder and chargeOrder stand in for your own factory and service code.)
The reply follows Stripe's documented error object: status 402 Request Failed, type: 'card_error', and a card_declined error code with a decline_code. stripe-node maps a 402 to StripeCardError, so your wrapper's real catch block runs against the real error class. If someone breaks the cents conversion, the first expect fails. If someone swallows the decline, the last one does.
The subtle part: mock payloads, not behavior
Mocking input data is fine. Mocking your code's behavior is not.
An auth fixture that hands your code a valid token payload is mocking input: it stands in for the identity provider at the boundary. That's correct. Reaching inside and stubbing your own permission check so it always returns true is mocking your code. Now the only thing under test is whether true === true.
Same with test data. Don't hardcode an order ID that only exists in one database on one machine. Create the record with a factory or pull a real one from your test database, so the test runs in anyone's environment. (I cover that step in the anatomy of an integration test, and keeping the database clean between tests in function vs class test database isolation.)
Why integration-first pays off
Once the boundary is the only thing you mock, integration tests become the natural default: a real request through routing, auth, service, wrapper and database, against an isolated database with real data. Unit tests still earn their place for pure utility functions used in many places, but they stop being the thing you reach for by reflex.
The payoff is trust. When the test goes green, the path a customer hits works. When it goes red, something real broke.
In a payments service the wrapper around Stripe is where dollars become cents and declines become order states, so a test that patches it out skips the code most likely to charge the wrong amount or lose a failed payment.
If you own the code, let it run; mock only the request that leaves your system.
The same rule in pytest
The idea doesn't depend on the stack. stripe-python's sync client uses requests by default, so responses can answer for Stripe the same way (use RESPX if your code calls Stripe through httpx):
tests/test_pay_order.pyimport responsesdef test_declined_card_marks_order_failed(client, order, db):with responses.RequestsMock() as stripe:stripe.post("https://api.stripe.com/v1/payment_intents",status=402,json={"error": {"type": "card_error", "code": "card_declined","decline_code": "insufficient_funds"}},)response = client.execute(pay_order_mutation, {"orderId": order.id})assert response.errors is Noneassert db.get_order(order.id).status == "payment_failed"
RequestsMock raises a ConnectionError for any request you didn't register, and by default fails the test if a registered one was never called, so it's strict in both directions.
When you're staring at a test wondering whether to patch something, ask one question: do I own this? If yes, let it run.