Debugging7 min read

Can't Reproduce a Production Bug Locally? Do This Instead

Author:Viraj Rakholiya

What is a Non-Reproducible Production Bug?

Definition Block: A non-reproducible production bug is an error, crash, or unexpected behavior that occurs in a live production environment but cannot be replicated in a developer's local testing environment or staging server. These bugs are typically driven by differences in data shape, concurrency, latency, or environment configuration rather than pure algorithmic flaws in the codebase. Because they depend on specific, often transient conditions unique to the production environment, they require an evidence-based debugging approach rather than traditional local reproduction attempts.

The "Works on My Machine" Dilemma

Every software engineer has been there: an alert fires, users are complaining, and the logs are bleeding red. You pull down the latest main branch, spin up your local environment, run the exact same steps, and... nothing. The app works perfectly.

A production bug you can't reproduce locally is almost never a simple logic bug. If it were a misplaced if statement or a basic syntax error, it would break everywhere. Instead, it is almost always an environment bug. It's about real data shapes (like null values your mock data never returns), real network latency (race conditions that your fast local database hides), actual environment variables, or complex user interaction sequences that are impossible to predict.

Stop trying to recreate it blind. When a bug refuses to show its face locally, the worst thing you can do is randomly tweak code hoping it might fix the unseen issue. You need to debug from captured production evidence instead.

Debug From Evidence, Not Reproduction

When local reproduction fails, your only source of truth is what actually happened in production. Here is a step-by-step reasoning on how to shift from a "reproduction" mindset to an "evidence-based" debugging mindset:

1. Pull the Exact Payload

Your first line of defense is robust, structured logging. A generic error message like "Failed to process order" is useless. You need the exact input that triggered the failure.

  • Implement Structured Logging: Ensure all logs output as JSON with a unique requestId that traces the journey of a single request across your microservices.
  • Log the Context: Always log the exception object, the request headers (sanitizing PII/auth tokens), and the specific payload parameters.
  • Example: If you're dealing with a nasty JavaScript error, understanding the exact object structure is crucial. For instance, read our guide on how to handle the classic TypeError: Cannot read properties to see why exact payload structures matter when things go wrong in production.

2. Watch the Replay

Sometimes the bug isn't in the payload; it's in the sequence of events. Users do things you would never think of—double-clicking submit buttons, opening multiple tabs sharing the same session state, or navigating backwards during a multi-step form.

  • Session Replay Tools: Tools like Sentry, PostHog, or LogRocket record the exact DOM mutations, network requests, and user clicks leading up to an error.
  • Identify State Corruption: Replays help you see the state of the application before it broke. This is invaluable for catching frontend race conditions or state management bugs in React/Vue applications.

3. Leverage AI and Autonomous Debugging with Relia

What if you didn't have to hunt down the logs, piece together the payload, and stare at session replays? This is where modern AI agents come in.

With Relia, the ultimate autonomous bug fixing tool, the failing request arrives as one comprehensive bundle. Relia captures the stack trace, the exact payload, the distributed trace, and the deploy SHA. But it goes a step further: it automatically analyzes the codebase and provides a proposed patch. There's nothing to reproduce manually because the evidence is the debugger.

Relia autonomously investigates the root cause by correlating the production environment data with your source code. If you're curious about the mechanics behind this, check out our deep dive on how autonomous bug fixing works. Relia handles the heavy lifting so your team can review the fix and ship it in minutes instead of hours. Visit relia.com or try it out at app.tryrelia.com.

The 5 Environment Differences to Check

If you must dig into the environment differences manually, focus on these five areas. They are the most common culprits for the "works locally, breaks in prod" phenomenon.

1. Data Shape and Quality

Your local database is likely seeded with perfect, idealized data. Production is messy. It has legacy rows from schema migrations 3 years ago, empty arrays where you expect populated objects, and unexpected null values.

  • Pro: Replicating real data uncovers edge cases quickly.
  • Con: Blindly copying production data locally is a massive security and privacy risk (PII). Always sanitize or selectively extract only the failing record's shape.

2. Network and Database Latency

Your local dev server has 0ms latency to your local PostgreSQL database. In production, the DB might be in another availability zone, or the connection pool might be saturated.

  • The Catch: Fast local services hide race conditions, optimistic UI failures, and timeout bugs.
  • The Fix: Introduce artificial latency to your local environment (e.g., using a proxy like Toxiproxy) to see how your app handles slow responses.

3. Environment Variables

A missing, misspelled, or stale environment variable in production is a classic silent killer.

  • The Check: Compare the exact keys (and the expected formats of the values) between your .env.local and your production deployment configuration.

4. Version Drift

Is your laptop running Node.js v20 while production is on v18? Are the database versions mismatched? Did a ^ in your package.json cause production to install a slightly newer, bugged patch version of a critical dependency?

  • The Fix: Use strict versioning (lockfiles), Dockerize your local environment to perfectly match the production OS and runtime, and enforce infrastructure-as-code.

5. Concurrency and Load

Locally, you are sending one request at a time. Production is handling hundreds of concurrent requests. This difference exposes thread-safety issues, database connection pool exhaustion, and memory leaks.

  • The Fix: Use load testing tools against a staging environment to simulate the high concurrency of production.

Isolate One Variable At A Time

When you're trying to track down the discrepancy, employ the scientific method. Do not change the data, the Node version, and the environment variables all at once.

  1. Change the Data Shape First: This is usually the easiest and most common culprit. Mock the exact API response or database row that failed.
  2. Alter Latency: If the data isn't the issue, add artificial delay to your local network requests.
  3. Verify Configuration: Triple-check environment variables and infrastructure settings.
  4. Simulate Load: Finally, if all else fails, use a tool like k6 to hammer your local or staging server to tease out concurrency bugs.

When the bug finally appears in exactly one of these configurations, you've isolated the root cause.

FAQ

Should I copy the production database locally?

Rarely the whole DB. Extracting the entire production database is a massive security risk and often violates data privacy laws (like GDPR or CCPA) due to PII. Instead, extract the specific failing record's shape, sanitize it, and replay that exact input. It is faster, safer, and far more targeted.

What if it's a race condition?

Replay with load testing tools (like k6 or JMeter) against a staging environment that mirrors production architecture. Single-request local testing will almost never show a race condition. You need concurrency to expose it.

How do I stop non-reproducible bugs long-term?

Validate your API boundaries rigorously using libraries like Zod, attach a unique requestId to every log across all services, keep your staging data shaped similarly to production, and adopt autonomous debugging tools like Relia to catch and patch issues before they escalate.

[ MORE ARTICLES ]

Read Next

View all →