Design API Gateway System

Design a production-grade API Gateway that provides authentication, authorization, rate limiting, request routing, and observability for multiple backend services.

Functional requirements

  • Terminate TLS and accept client requests for multiple backend APIs.
  • Authenticate requests (JWT/OAuth/API keys) and attach identity to the request context.
  • Authorize requests (RBAC/ABAC) based on route, method, tenant, and user.
  • Route requests to the correct internal service based on host/path/headers.
  • Apply per-tenant quotas and per-user/IP rate limits.
  • Emit access logs, metrics, and traces with correlation IDs.
  • Support safe configuration updates (new routes, policy changes) with rollback.

Non-functional requirements

  • p95 gateway added latency should be under 15ms; p99 under 40ms under normal load.
  • 99.9% availability for the gateway tier (multi-AZ; no single point of failure).
  • Graceful degradation under overload (shed best-effort work before failing core routing/auth).
  • Strong security posture: input validation, WAF rules, least-privilege, secrets hygiene.
  • High observability: debug a single request end-to-end within minutes (trace + logs + metrics).
  • Config changes should propagate safely within minutes with auditability and rollback.

How the design evolves

Stage 1: Edge gateway with auth + routing

Understand the basics of an API gateway: TLS termination, authentication, and routing to internal services.

What was missing: Clients called each backend service directly, so TLS, authentication and routing logic were duplicated in every service and inconsistently enforced.

Why that's risky: Only ~8% of requests reach the auth service. If the cache cluster is lost, auth load jumps 12x instantly; the auth tier needs surge headroom or the gateway must fail closed.

What gets added: An edge load balancer (TLS + WAF) in front of a 10-replica stateless gateway fleet, plus an auth/policy cache so token validation does not hit the auth service on every request.

Trade-offs: The gateway becomes a shared dependency and a place logic tends to accumulate. Keep it to cross-cutting concerns only — routing, authn/authz, limits, observability — and never business logic.

Stage 2: Traffic control: quotas, rate limits, and safe rollouts

Implement traffic control features like per-tenant quotas and rate limits to protect backend services, and design a safe rollout strategy for gateway config changes.

What was missing: Any single tenant could saturate the gateway and starve everyone else, and every routing or policy change required a redeploy with no safe rollback.

Why that's risky: Rate-limit counters are the classic hot shard. Partitioning by tenant+api-key avoids one noisy tenant serialising the whole limiter.

What gets added: A distributed token-bucket rate limiter sharded by tenant/api-key ahead of the gateway, a versioned route/policy config store, and a canary deployment of the Order Service taking 10% of order traffic.

Trade-offs: Rate limiting adds ~2.5ms and a shared-state dependency. It fails open so a Redis shard loss degrades fairness rather than availability — the right trade for an availability-first edge, but it means limits are best-effort during incidents.

Stage 3: Production hardening: rate limits, observability, async logs

Implement rate limiting, observability, and asynchronous logging to ensure the gateway can handle production traffic safely and efficiently.

What was missing: There was no way to debug a single request end-to-end, prove SLO compliance, or retain access logs — and doing any of that synchronously would have blown the latency budget.

Why that's risky: If the log pipeline were synchronous and critical, a slow analytics store would add latency to every API call. Marking these edges async and non-critical is what keeps that blast radius contained.

What gets added: A fire-and-forget access-log pipeline (Kafka buffer → 20 ingester workers → analytics store) and an OTLP metrics/tracing collector fed by 5% head-based sampling.

Trade-offs: Async logging means logs can lag or, in a total buffer loss, be dropped. That is acceptable for observability data and unacceptable for anything billing-related, which must stay synchronous.

Frequently asked questions

JWT validation vs token introspection — which should the gateway use?

JWT validation is fast and avoids an auth dependency in the hot path (good for p99). Introspection enables immediate revocation and central control but adds latency and creates a dependency. Many production systems use JWT validation with short token TTLs plus a revocation mechanism for high-risk cases.

How do you do rate limiting without making the limiter a bottleneck?

Scope limits correctly (per-tenant/user/key), keep checks O(1), and avoid a single centralized counter. Common approaches: sharded counters in Redis, local leaky-bucket with periodic sync, or tiered limiting (edge coarse limit + gateway fine limit).

Where should timeouts/retries live — gateway or services?

Both, but with clear budgets. The gateway should enforce an overall deadline and avoid unbounded retries. Services often need their own timeouts/circuit breakers for downstream dependencies. The key is preventing retry storms and respecting a per-request latency budget.

PRISM
System Design Interview
Round 1 of 4 · Architecture Design · 60:00 remaining
PRISM logo
AI Interview
Interview Prep
Interview Challenges
Design Your Own System NEW
System Architectures
Interactive Roadmap System Design Guides
Notifications
  • No new notifications
Feedback
Signed in
Phase
Design Your Own System
Phase 01: Thinking in Systems
Upcoming

Components

User
CDN
Load Balancer
Server
Cache
Database
Blob Storage
Search Index
Queue
Worker
Rate Limiter
Service
API Gateway
Reverse Proxy
WebSocket Server
Third-party API

Inspector

Notes
Use clear names so your design intent is easy to understand.
Good
Name by business meaning
"Order API", "Restaurant Service"
Avoid
Generic names = zero signal
"Server 1", "API", "Queue"
A short description for each component makes feedback much better.
100%
Start by identifying:
  • Users & entry points
  • APIs & services
  • Databases & storage
  • Traffic flow & scale

Round 1 of 4 Architecture Design

Run a simulation to see results.

Time Remaining
60:00
System Design Interview

What are the core functional requirements?
What are the key non-functional constraints?

Questions

Start Evaluation to unlock questions.

Components Added
No components listed
Click "+ Add" to document components introduced in this stage.
Design Decisions

Click Simulate to run your design and see results here.

Internal notes — not shown to learners.

EVALUATE MODE

Test yourself like it's the real thing.

A structured 4-module evaluation that mirrors how top companies assess system design candidates.

Architecture Design
Draw your system on the canvas. Define components, connections, and data flow.
FR & NFR Requirements
Answer functional and non-functional requirement questions about your design.
MCQ Round
Multiple choice questions testing your depth on the chosen system.
Tradeoff Analysis
Justify your design decisions and defend your architectural tradeoffs.
AI Report Generated
A R S
Used by engineers preparing for FAANG & top-tier companies
Choose a Problem
No problem selected
  • 30 min
  • 45 min
  • 60 min
Round 2 of 4
MCQ Round
Answer multiple-choice questions based on your design.

Exit Interview?

You're in the middle of an interview session. Leaving now will end your current attempt.

Your progress will be saved.

Open a saved design

Select a design to load into the canvas.

My Evaluations

Your past evaluation sessions

Here’s a simple request flow that follows the expected layer order.

External User
→
Edge CDN → API Gateway → Load Balancer
→
Compute App Servers / Services
→
DataAccess Cache
→
Storage Database / Search Index
→
Async Queue → Worker

Tip: keep arrows moving forward through layers (Edge → Compute → Storage). Avoid sending storage back to compute.

Evaluation Instructions

Read the rules carefully before starting. The test auto-submits on refresh.

Before you start

  • Build your architecture on the canvas. The timer starts when you click Start Evaluation.
  • Don't forget to answer Functional Requirement and Non Functional Requirements.
  • When satisfied with your design, click Next to lock it and view the questions.
  • Please answer final step questions to complete the evaluation.

Dos

  • Do read each question carefully before answering.
  • Do include required components to maximize component coverage.
  • Do save a copy of your design if you want to keep it before submission.

Don'ts

  • Don't refresh or close the tab during an active evaluation — this will auto-submit your answers.
  • Don't switch app modes or open another tab while the evaluation is running.
  • Don't attempt to edit the design after clicking Next; the workspace will be locked.

All the best!!

Confirm

Input

Notice

Evaluation Report:

Evaluation Complete

Generating Your Report

Hang tight — our AI is evaluating your design…

Did you know?

Loading…

Share feedback

Tell us what worked well and what we can improve.

Let's personalize this

Answer a couple of quick questions so we can tailor your journey and missions.

You can change this anytime from your Profile.

Your personalized missions are ready

We tailored these first steps based on your answers.

    PRISM Welcome Gift

    This is a personal welcome gift from PRISM.

    Congratulations.

    You explored PRISM.

    You earned Apprentice.

    As a welcome gift, unlock Full PRISM Access for the configured trial duration.

    This starts Trial. Trial timer begins only after you activate this gift.

    Welcome to PRISM

    We've prepared a personalized Apprentice Journey based on your goals and experience.

    This journey introduces you to the capabilities of PRISM that are most relevant to you.

    Complete all 8 missions to earn your Apprentice title. 8 MISSIONS

    PRISM Surprise Offer

    Complete your Apprentice Journey to unlock a special gift from PRISM.

    • No payment required
    • No credit card required
    • Just complete the journey
    MISSION CONTROL
    0 / 8 missions complete
    NEXT UP Continue your missions
    View full roadmap →
    Mission Complete 0 / 8 Completed Next: Keep going
    SYSTEM BRIEF

    ⬤ System Constraints

    What the system must do — every item is a user-facing behaviour your architecture must support.

      ↑ Engineering Constraints

      These are the failure modes you must design against — latency SLAs, durability targets, traffic ceilings.

        ⇆ Architecture Constraints

        ◈ Core Concepts to Master

        Your Journey
        PHASE – –
        0 / 0 0%
        0
        Mock Interview Checklist

        Pick a topic to start

        Explore concept overviews, real-system examples, key tradeoffs, and interview talking points for each roadmap section.

        Topic-Wise Progress
        Experience Points 0 XP
        Read subtopics & solve challenges to earn XP
        Theory Read +0 XP
        Solved +0 XP
        Streak Bonus +0 XP
        Theory Read 0%
        — Mastered — Solved
        Weekly Streak 0 day streak
        Mon
        Tue
        Wed
        Thu
        Fri
        Sat
        Sun
        Keep going — log in daily to build your streak!
        0 0%
        Skill Profile
        Recommended Next
        🎯 Your Focus

        You haven't explored enough yet.

        Focus on
        → Understanding System Design
        → Estimating Scale
        Next Action
        Continue → Next: –
        Mock Interview Checklist
        Architecture DNA
        Engineering Profile
        Phase Mastered

        You've conquered this phase. These are the skills you now own:

          +500 XP

          Engineering Profile

          Company Interview Paths

          Progress Summary