Flask guide · Python SDK

Flask rate limiting,keyed to your customers.

Wrap your Flask app once. Every request is checked against the caller’s API key, charged against their credits, and held to their plan’s rate limit before your view runs.

  • Keys issued per customer
  • Credits that refill
  • Limits per plan, not per process

Official SDKs with drop-in middleware for the stack you already run

  • Python
  • Node.js
  • Go
  • Rust
  • PHP
  • .NET
  • Java
API requests validated
100M+
Average key validation
<5ms
Average analytics ingest
<5ms
Check and log, end to end
<10ms

The usual Flask rate limiter counts requests. It doesn’t know your customers.

Flask-Limiter is the standard answer: default limits for the app, a decorator for route-specific ones, and a storage URI that decides where counts live.

With Flask-Limitertoday
# app.py — with Flask-Limiterfrom flask import Flask
from flask_limiter import Limiter
from flask_limiter.util import get_remote_address

app = Flask(__name__)
limiter = Limiter(
    get_remote_address,
    app=app,
    default_limits=["200 per day", "50 per hour"],
    storage_uri="memory://",
)

@app.route("/search")
@limiter.limit("10 per minute")
def search():
    return {"results": []}
  • Keyed by IP address by default, not by the customer who is paying
  • memory:// storage is per process — Gunicorn workers each count separately
  • Needs Redis (or another shared store) before the limit is correct on more than one server
  • No API keys: issuing, hashing, scoping, and revoking them is still yours to build
  • No credit balance: a request can’t cost 5 on one route and 1 on another, or refill each month
  • One limit for everyone — no per-plan limits for Free, Pro, and Enterprise customers
With ReqKey in Flaskone middleware
  • API keys issued per customer, prefixed, hashed, and revocable from the dashboard
  • A credit balance per customer — price each route, refill every hour, day, week, or month
  • Rate limits set per plan, shared by all of a customer’s keys, from 1 second to 24 hours
  • The same count on every worker, instance, and region — no Redis to run
  • Over the limit? A 429 with Retry-After, and no credits charged
  • Every request logged per customer — status, latency, endpoint — in under 5ms
Your view reads the verdict from request.environ["reqkey.decision"]. See the code

Flask in three steps, one of them code.

ReqKey runs inside your API, not in front of it. Your server asks one question per request and gets an answer in under 5ms.

  1. 1

    Install the SDK

    pip install "reqkey[flask]"

  2. 2

    Set your project key

    Copy it from the dashboard into REQKEY_PROJECT_KEY. It stays on your server.

  3. 3

    Add the Flask middleware

    Every request is checked, charged, and logged before your handler runs.

app.pypip install "reqkey[flask]"
import os
 
from flask import Flask, request
from reqkey.flask import ReqKeyMiddleware
 
app = Flask(__name__)
 
# Flask stays the same object — only its wsgi_app callable is wrapped
app.wsgi_app = ReqKeyMiddleware(
app.wsgi_app,
project_key=os.environ["REQKEY_PROJECT_KEY"],
api_id="api_payments",
mode="both",
key_name="X-MyStartup-Key",
exclude_paths=("/health", "/static/*"),
)
 
 
@app.post("/payments")
def create_payment():
decision = request.environ["reqkey.decision"]
return {"created": True, "credits_remaining": decision.credits_remaining}
Also in Python: FastAPI, DjangoFull Python reference

Every request gets one of these answers

  • 200Valid key, credits charged — your handler runs
  • 402Out of credits
  • 403Key disabled, or not allowed on this API
  • 429Over the rate limit — no credits charged

Where ReqKey sits

Your customer
Your APIReqKey, under 5ms
Your handler

Responses go straight back to your customer. ReqKey sees the key check and the log line, nothing else.

Credits

Charge each route what it costs you.

A lookup can cost 1 credit and a render 5, from the same balance. Excluded paths are never validated, charged, or recorded. In Python:

Per-route credits · Python
# Exact paths or trailing-* prefixes: never validated, charged, or recorded
exclude_paths=("/health", "/openapi.json", "/docs/*", "/cron/*")
 
# Or decide per request with a sync or async resolver
should_protect=lambda request: request.url.path.startswith("/api/")
 
# Charge different endpoints differently
def credits_for(request):
if request.method == "POST" and request.url.path == "/images":
return 5
return 1
 
app.add_middleware(ReqKeyMiddleware, ..., credits=credits_for)

Rate limits

Set limits on plans, not in code.

Your Flask code never hard-codes a number. Each plan carries its credits, refill, and rate limit; moving a customer to Pro changes all three with no deploy.

PlanCreditsRate limit
Free1,000 / month5 req / s
Pro50,000 / month50 req / s
Scale1,000,000 / month500 req / s
429Over the limit, a call is answered with Retry-After and costs nothing — no credits and no quota.

Example plans. You name them and pick the numbers.

Notes for Flask teams hit in production.

Your Flask object stays the same

Only app.wsgi_app is wrapped, so blueprints, error handlers, and extensions keep working untouched.

Gunicorn and friends

Because the count lives in ReqKey, the limit holds whether you run one Gunicorn worker, twelve, or several machines.

Any WSGI app

The same middleware works for Bottle, Pyramid, or anything that speaks PEP 3333 through the generic WSGI adapter.

Flask rate limiting: the questions teams ask.

Something else? Ask the team or read the docs.

  • You can, for example to throttle anonymous traffic by IP on public pages. For customer traffic, ReqKey already applies each customer’s plan limit, so you don’t need both on the same routes.

  • Pass exclude_paths such as ("/health", "/static/*"). Excluded paths are never validated, charged, or recorded.

  • On their plan or on the customer (consumer) in ReqKey — a number of requests per window from 1 second to 24 hours, shared by all of that customer’s keys. Your Flask code never hard-codes a limit, so upgrading a customer is a dashboard change, not a deploy.

  • A check averages under 5ms. ReqKey runs inside your app as middleware, not as a gateway in front of it, so responses go straight back to your customer.

  • You choose: fail closed and answer 503, or fail open and let requests through. Invalid keys are denied either way, and a validation is never retried, so no one is charged twice.

Ship API keys in Flaskin five minutes.

Free for your first 5 million requests every month. No card, no gateway, no rewrite.