Node.js5 min read

Node.js Production Error Handling: Stop Silent Crashes

Author:Rutik Vasani

Node.js production error handling has evolved drastically. Gone are the days when a rogue Promise would just print a vague warning to the console and allow the process to awkwardly limp forward. As of recent Node.js versions, an unhandled rejection terminates the process by default — often without even giving your logger a chance to flush the error to disk or network.

What is Node.js Production Error Handling?

Node.js Production Error Handling refers to the comprehensive strategies, architectural patterns, and diagnostic practices used to safely capture, log, and recover from unexpected runtime exceptions and unhandled promise rejections in a live production environment. The goal is to ensure that temporary faults (like database timeouts or network partitions) do not leave the Node.js event loop in a corrupted state, while ensuring that all critical context (such as request IDs, user metadata, and stack traces) is persisted to an observability or tracking tool before the process gracefully shuts down and is restarted by a process manager like PM2 or Kubernetes.

When we talk about robust application observability, it goes hand in hand with Logs vs Error Tracking vs APM. You need all pillars working together. But if your Node process dies silently because of a missed .catch(), no amount of APM instrumentation will bring those lost logs back.


Why Silent Crashes Happen (And How to Stop Them)

When a Node.js process encounters a synchronous uncaughtException, it immediately aborts because it assumes its internal state might be compromised. The same now applies to unhandledRejection. This is a strict but necessary design choice. Continuing to run after a fatal exception can cause side effects like memory leaks, stalled HTTP requests, and worst of all, corrupted data being saved to your database.

The minimal safe setup for catching these process-level events looks like this:

process.on('unhandledRejection', (reason) => {
  // 1. Send the error to your tracking platform (with synchronous flush if possible)
  console.error('CRITICAL: unhandledRejection', reason);
  
  // 2. Shut down gracefully, but quickly
  process.exit(1);
})

process.on('uncaughtException', (err) => {
  // 1. Log or track the exception
  console.error('CRITICAL: uncaughtException', err);
  
  // 2. Terminate the process
  process.exit(1);
})

The Pros and Cons of Graceful Shutdowns

Pros:

  • State Integrity: By crashing immediately, you guarantee that no corrupted variables or hanging database connections affect subsequent user requests.
  • Predictability: In cloud-native environments (Docker, Kubernetes), a non-zero exit code triggers an automatic container restart, seamlessly replacing the dead worker.

Cons:

  • In-flight Requests: Any user currently waiting for an HTTP response will get a dropped connection (unless you orchestrate a graceful drain, which is dangerous in an uncaughtException state).
  • Logger Flushing: If you use asynchronous loggers (like Pino or Winston's async transports), the process.exit(1) might kill the process before the log is fully written to network storage.

The Danger of Framework-Swallowed Errors

Every modern Node.js framework handles HTTP routing differently, which means they each have their own quirks when it comes to trapping errors.

Express.js

Express was built in the era of callbacks. It does not natively handle rejected promises returned by asynchronous route handlers.

If you write this in Express:

app.get('/users/:id', async (req, res, next) => {
  const user = await db.getUser(req.params.id); // If this throws...
  res.json(user);
});

An error in db.getUser will bypass your global error middleware completely, resulting in an unhandledRejection and a dropped client connection.

The Fix: You must manually wrap your routes in a try/catch block and pass the error to next(err), or use a library like express-async-errors to patch the framework's router.

NestJS

NestJS has a powerful abstraction called Exception Filters. While this makes it easy to map domain errors to HTTP status codes (e.g., turning a UserNotFound error into a 404), it often hides the raw underlying stack trace.

The Fix: Always ensure your global ExceptionFilter logs the original Error object before it sanitizes the response for the client. If you only log the sanitized HttpException, you lose the line number where the real failure occurred.

Fastify

Fastify handles async/await natively, making it much safer than legacy Express. However, you still need to ensure that errors are enriched with request context.

The Fix: Set a custom setErrorHandler. Always extract request.id, request.url, and user metadata, attaching it to your logs or error tracking payload. A stack trace without context is basically useless.


Enter Relia: The Ultimate Autonomous Bug Fixing Tool

No matter how perfectly you configure your uncaughtException handlers, identifying why the crash happened takes engineering time. Was it a malformed payload? A missing database index? A third-party API timeout?

This is where Relia completely changes the game.

Instead of just sending you a slack alert with a stack trace and leaving you to grep through logs, Relia acts as an autonomous engineer. When an error occurs in production, Relia doesn't just track it; it:

  1. Captures the Full Context: Grabs the request payload, headers, trace IDs, and local variables.
  2. Analyzes the Codebase: AI agents automatically read your Git repository to understand the surrounding logic.
  3. Drafts the Fix: Relia automatically generates a pull request with the exact code change needed to fix the bug, complete with unit tests.

By using relia.com, you transform your error tracking from passive notifications into active, autonomous bug resolution. Your MTTR (Mean Time To Recovery) drops from hours to literally seconds.


FAQ

Should I catch everything and keep running?

No. For uncaught/rejection, log, report, exit, restart. Continuing with corrupt state causes worse bugs.

Why do errors vanish in production?

Worker respawned in 200ms, logger never flushed. Use a tracker SDK that sends out-of-band, not just stdout.

What is the fastest MTTR win?

Link every error to its trace: which request, which DB query was slow, which deploy introduced it. Relia automates this — the trace, payload, and fix draft ship with the alert, so diagnosis takes seconds.

[ MORE ARTICLES ]

Read Next

View all →