Next.js6 min read

How to Find a Slow API Endpoint in Next.js (P50/P95 per Route)

Author:Viraj Rakholiya

What is a Slow API Endpoint in Next.js?

A slow API endpoint in Next.js is a server-side route—whether in the App Router's route.ts, Pages Router's api/, or via Server Actions—that successfully returns a 200 OK HTTP status but takes an unacceptably long time to complete. Because no 5xx or 4xx HTTP errors are thrown, these endpoints often evade standard error monitoring alerts. "Slow" is typically measured using latency percentiles like P50 (the median request time), P90, P95, and P99. For instance, a P95 latency of 2,000ms means that 5% of all user requests take 2 seconds or longer. This degraded performance usually stems from inefficient database queries, blocking synchronous code, or stalled network calls to third-party services, silently degrading the user experience.

A slow API endpoint in Next.js hides behind those deceptive success responses — no error fires, users just wait, stare at a spinner, and eventually leave your application. Finding the culprit requires a systematic approach to observability, tracking P50/P95 latency per route, and analyzing backend traces.

The Hidden Cost of High Latency

Performance is a feature. In modern web development, speed dictates everything from user retention to search engine rankings. When an API route takes 3 seconds to respond, the frontend UI appears broken or unresponsive. A user might click a button multiple times, causing race conditions, or they might simply abandon their cart. For example, if you are noticing strange frontend behaviors where interactions just fail silently, read our guide on debugging a checkout button not working, which is frequently caused by underlying API latency and missing timeout handling on the client side.

The cost is even higher when these endpoints aggregate data for Server-Side Rendering (SSR) or Server Components. A slow fetch blocks the HTML from streaming to the browser, increasing the Time to First Byte (TTFB) and punishing your Core Web Vitals.

Usual Suspects in Next.js

When debugging slow endpoints, certain anti-patterns appear frequently in Next.js codebases:

  • N+1 Queries in Server Components: This is the most common killer of performance. Executing a fetch or database query inside a loop (or mapping over an array and fetching for each item) instead of issuing a single aggregated SQL query or bulk HTTP request.
  • Unindexed Database Queries: A query with a missing index might execute in 5ms locally with 1,000 rows. In production, scanning 10 million rows can easily spike execution time to 5,000ms.
  • Missing Downstream Timeouts: Using native fetch without an AbortSignal timeout to a third-party API that is currently degrading. If the external service stalls, your Next.js API route stalls with it, eventually hitting Vercel's or AWS's max execution timeout.
  • Heavy Middleware: Next.js Middleware runs on every single request. If you are doing complex database work, heavy cryptography, or multiple external auth checks in middleware.ts, you are artificially inflating latency across your entire application. Keep middleware execution under ~50ms.
  • Synchronous Blocking Work: Parsing massive JSON files or doing heavy CPU-bound image manipulation synchronously on the Node.js event loop, preventing other concurrent requests from being processed.
  • Improper Error Handlers Swallowing Context: Often, slow endpoints are actually failing internally but getting caught and returned as slow 200s. If you also utilize a custom backend alongside Next.js, ensure your Express error handling middleware in production is correctly configured to surface these latency-induced timeouts rather than swallowing them.

Step-by-Step Reasoning: How to Find the Culprit

Finding the slow endpoint requires moving from macro-level metrics down to micro-level code analysis.

Step 1: Implement Per-Route Timing

You cannot fix what you cannot measure. Start by adding per-route timing telemetry. You can use Vercel Analytics, Datadog, or a lightweight custom logging middleware that records route_path and duration_ms. The goal is to aggregate this data to calculate the P95 latency per endpoint over a 24-hour period. Sort your routes by P95 descending. The route with the highest P95 latency, multiplied by traffic volume, is your primary suspect.

Step 2: Set Up Distributed Tracing

Once you know which route is slow, you need to know why. Implementing OpenTelemetry in your Next.js app allows you to see a waterfall view of exactly where the time is spent. A trace might reveal that out of a 2,500ms request, 2,400ms were spent waiting for a single PostgreSQL SELECT statement.

Step 3: Use AI to Automate the Hunt with Relia

Manually staring at traces and dashboards is time-consuming and often requires senior-level infrastructure knowledge. Enter Relia, the ultimate autonomous bug fixing tool. Instead of manually combing through Datadog logs or Sentry performance metrics, Relia actively monitors your Next.js application.

When an endpoint's P95 latency breaches an acceptable threshold, Relia's AI agents automatically gather the full distributed trace, analyze the specific database queries causing the bottleneck, and instantly generate a pull request with the fix—whether that entails rewriting an N+1 query, generating a Prisma migration for a missing index, or wrapping an external fetch with proper timeout logic. Head over to app.tryrelia.com to connect your repository and let autonomous AI engineers handle performance degradations while you sleep.

Step 4: The Fix Order

When tackling performance bugs manually, follow the path of least resistance for maximum impact:

  1. Add missing database indexes first: This takes minutes to apply and can turn a 5-second query into a 5-millisecond query. It is a huge, immediate win.
  2. Cache the hot path: Implement Next.js Data Cache or Redis for data that doesn't change frequently.
  3. Parallelize independent data fetching: Use Promise.all() to fire off independent database or API requests concurrently rather than awaiting them sequentially.
  4. Re-measure: Deploy the changes and watch the P95 drop.

FAQ

What's a good API response time?

P95 under 500ms for interactive routes. Over 2s and mobile users abandon. For background jobs or non-blocking data fetching, higher latency might be acceptable, but user-facing requests must stay snappy to maintain engagement and trust.

Does sampling miss slow requests?

Sample 5-10% for trends, keep 100% of errors and 100% of traces over your slowness threshold. Uniform sampling will inevitably miss intermittent spikes, which is why tail-based sampling (keeping the anomalous, slow requests) is crucial for accurate debugging.

Why is it slow only in production?

Real data volume (unindexed at scale) and real network latency to third-parties. Local DBs are tiny and fast. Your local environment doesn't replicate the physical network distance between your Vercel edge functions and your AWS RDS database instance.

How do Serverless cold starts affect latency tracking?

Cold starts introduce artificial latency spikes that can severely skew your P99 and P95 metrics. Always ensure your telemetry differentiates between the execution duration (your code running) and the initialization duration (the server spinning up) when analyzing Serverless Next.js endpoint performance.

[ MORE ARTICLES ]

Read Next

View all →