Why Rate Limiting Matters

I've learned the hard way that REST API design without rate limiting is a ticking time bomb. During my early days at CodeBrew Labs, we shipped a public API without proper protection. Within weeks, a client's misconfigured mobile app was hammering our endpoints with thousands of requests per second. Our database went down. Our ops team was paged at 2 AM. It was chaos.

That experience taught me that rate limiting isn't just about preventing abuse—it's about protecting your infrastructure, ensuring fair resource allocation, and maintaining API performance for legitimate users. Whether you're building a SaaS platform, a mobile backend, or a third-party integration layer, rate limiting is non-negotiable.

I've since implemented rate limiting across multiple production systems, and I want to share the practical strategies that actually work.

Understanding Rate Limiting Algorithms

Before jumping into code, let's cover the three main approaches I use in production.

1. Token Bucket Algorithm

This is my go-to for most scenarios. Imagine a bucket that fills with tokens at a fixed rate. Each request consumes one token. If the bucket is empty, the request is rejected.

  • Pros: Handles bursts gracefully, fair, easy to understand
  • Cons: Slightly more complex than other methods
  • Best for: Public APIs, multi-tier pricing

2. Sliding Window Log

Track individual request timestamps in a rolling window. If you exceed your limit in the last N seconds, reject the request.

  • Pros: Very accurate, no bursts allowed
  • Cons: Memory-intensive at scale
  • Best for: Strict rate enforcement, webhooks

3. Fixed Window Counter

Simple: reset a counter every N seconds. If it exceeds the limit, reject requests for that window.

  • Pros: Minimal overhead, easy to implement
  • Cons: Vulnerable to burst attacks at window boundaries
  • Best for: Internal APIs, non-critical endpoints

In production, I usually combine token bucket for user-facing APIs and fixed window for internal services.

Implementing Rate Limiting in Node.js

Let me walk you through a practical Node.js implementation using Express and Redis. I've used this exact pattern across multiple backends with great results.

import express from 'express';
import redis from 'redis';
import { promisify } from 'util';

const app = express();
const redisClient = redis.createClient();
const getAsync = promisify(redisClient.get).bind(redisClient);
const incrAsync = promisify(redisClient.incr).bind(redisClient);
const expireAsync = promisify(redisClient.expire).bind(redisClient);

// Token Bucket Rate Limiter Middleware
const rateLimitMiddleware = (limit = 100, windowInSeconds = 60) => {
  return async (req, res, next) => {
    const userId = req.user?.id || req.ip;
    const key = `rate_limit:${userId}`;
    const ttl = windowInSeconds;

    try {
      const current = await incrAsync(key);
      
      // Set expiration on first request in window
      if (current === 1) {
        await expireAsync(key, ttl);
      }

      // Set rate limit headers
      res.setHeader('X-RateLimit-Limit', limit);
      res.setHeader('X-RateLimit-Remaining', Math.max(0, limit - current));
      res.setHeader('X-RateLimit-Reset', Math.floor(Date.now() / 1000) + ttl);

      if (current > limit) {
        return res.status(429).json({
          error: 'Too Many Requests',
          message: 'Rate limit exceeded. Please try again later.',
          retryAfter: ttl
        });
      }

      next();
    } catch (err) {
      // If Redis fails, fail open (allow request)
      console.error('Rate limiter error:', err);
      next();
    }
  };
};

// Apply different limits to different endpoints
app.get('/api/data', rateLimitMiddleware(1000, 60), (req, res) => {
  res.json({ data: 'public endpoint' });
});

app.post('/api/auth/login', rateLimitMiddleware(10, 60), (req, res) => {
  res.json({ token: 'auth_token' });
});

app.listen(3000);

This implementation uses Redis for distributed state, which is critical when running Node.js across multiple servers. The counter increments atomically, and TTL is auto-managed.

💡 Pro Tip

Notice how I set X-RateLimit-* headers. This lets clients know their remaining quota without making extra requests. Always do this in production REST APIs.

Implementing Rate Limiting in Laravel

Laravel makes this incredibly easy with built-in middleware, but I often need custom logic for specific use cases. Here's how I handle it:

<?php

namespace App\Http\Middleware;

use Closure;
use Illuminate\Cache\RateLimiter;
use Illuminate\Http\Request;

class CustomRateLimit
{
    protected RateLimiter $limiter;

    public function __construct(RateLimiter $limiter)
    {
        $this->limiter = $limiter;
    }

    public function handle(Request $request, Closure $next)
    {
        $key = 'api_limit:' . ($request->user()?->id ?? $request->ip());
        $limit = 100;
        $decayMinutes = 1;

        if ($this->limiter->tooManyAttempts($key, $limit)) {
            $retryAfter = $this->limiter->availableIn($key);
            
            return response()->json([
                'error' => 'Too Many Requests',
                'message' => 'Rate limit exceeded.',
                'retry_after' => $retryAfter
            ], 429)->header('Retry-After', $retryAfter);
        }

        $this->limiter->hit($key, $decayMinutes * 60);

        $response = $next($request);

        return $response
            ->header('X-RateLimit-Limit', $limit)
            ->header('X-RateLimit-Remaining', $this->limiter->remaining($key, $limit))
            ->header('X-RateLimit-Reset', now()->addMinutes($decayMinutes)->timestamp);
    }
}

// In your routes/api.php
Route::middleware('throttle:custom-rate-limit')->group(function () {
    Route::get('/api/data', [DataController::class, 'index']);
    Route::post('/api/auth/login', [AuthController::class, 'login']);
});

Laravel's cache layer handles the complexity. By default it uses Redis if configured, and automatically manages TTLs and atomic increments.

⚠️ Configuration Note

Make sure your Laravel .env has CACHE_DRIVER=redis. File-based caching won't work for distributed rate limiting across multiple servers.

Redis-Based Approach for Distributed Systems

At Raybit, we handle high-traffic APIs. A simple in-memory limiter won't work when you have requests hitting different servers. Redis is the answer.

Here's why I always reach for Redis:

  • Atomic operations (INCR is thread-safe)
  • Automatic key expiration (TTL management)
  • Sub-millisecond performance
  • Works across distributed servers
  • Easy integration with both Node.js and Laravel

For Node.js, I use redis or ioredis directly. For Laravel, just set CACHE_DRIVER=redis and the framework handles it.

"We went from 5-second API response times during traffic spikes to consistent 200ms responses after implementing Redis-based rate limiting. It forced us to be smart about request queuing."

One optimization I've found: use Redis Lua scripts for complex rate limiting logic. This ensures atomicity without round-trips:

const rateLimitScript = `
local key = KEYS[1]
local limit = tonumber(ARGV[1])
local ttl = tonumber(ARGV[2])

local current = redis.call('incr', key)
if current == 1 then
  redis.call('expire', key, ttl)
end

if current > limit then
  return {0, current, ttl}
else
  return {1, current, ttl}
end
`;

const result = await redisClient.eval(
  rateLimitScript,
  1,
  `rate_limit:${userId}`,
  100,
  60
);

const [allowed, current, ttl] = result;
if (!allowed) {
  res.status(429).json({ error: 'Rate limit exceeded' });
}

Practical Considerations & Gotchas

Different Limits for Different User Tiers

One of the first things I implement in production APIs is tiered rate limiting. Free users get 100 req/min, pro users get 10,000 req/min:

const getTierLimit = (user) => {
  const tiers = {
    free: { requests: 100, window: 60 },
    pro: { requests: 10000, window: 60 },
    enterprise: { requests: 100000, window: 60 }
  };
  return tiers[user.tier] || tiers.free;
};

const rateLimitMiddleware = async (req, res, next) => {
  const tier = getTierLimit(req.user);
  const key = `rate_limit:${req.user.id}`;
  
  const current = await incrAsync(key);
  if (current === 1) await expireAsync(key, tier.window);
  
  if (current > tier.requests) {
    return res.status(429).json({ error: 'Rate limit exceeded for your tier' });
  }
  next();
};

Handling Redis Failures

What happens if Redis goes down? Always fail open. Don't let rate limiting become a single point of failure. In my code examples above, I wrap Redis calls in try-catch and let requests through if Redis fails.

Handling Client Retries

Return proper HTTP status codes and headers:

  • 429 Too Many Requests — Always use this status
  • Retry-After header — Tell clients when to retry
  • X-RateLimit-* headers — Let clients track their quota

Monitoring & Alerting

In production, I track how many requests hit the rate limit. A spike in 429s often indicates a bug in client code or a coordinated attack. Set up monitoring on your 429 response rate.

Key Takeaways

  • Rate limiting is essential for API performance and security. Use token bucket for most cases, implement it as middleware, and always use Redis for distributed systems.
  • Node.js and Laravel both have excellent rate limiting support. Node.js needs custom middleware; Laravel has built-in throttle middleware that works out of the box.
  • Always return proper HTTP headers (X-RateLimit-*, Retry-After) so clients can handle limits gracefully without wasting requests.
  • Fail open on infrastructure failures—don't let rate limiting become a bottleneck. If Redis is down, let requests through and trust your database to handle it.
  • Implement tiered limits based on user subscription level. This lets you monetize better and protects free-tier users from being abused.