Documentation

Revl Gateway documentation

Revl Gateway is an open-source, bring-your-own-key gateway that sits between your application and the Anthropic and OpenAI APIs and adds exact-match caching, PII and secret masking, retries and per-tenant rate limits.

Status. The hosted gateway at https://api.proxyrevlvay.com is invite-only and is being rolled out to the waitlist. It is not generally available. You can self-host the open-source core on your own Cloudflare account today. This page documents core v0.1.

On this page

How BYOK works

BYOK means bring your own key. The gateway holds no model-provider credentials. Every request carries two keys, and they do different jobs.

Your app

Sends both keys as request headers.

provider key Revl key

Revl Gateway

Checks the Revl key, then drops it. Never stores or logs the provider key.

provider key no Revl key

Provider

Anthropic or OpenAI, under your account, your billing and their terms.

Two keys travel with every request. Only your provider key reaches the provider, and only for that request.
  • Your provider API key travels in that provider's standard header: x-api-key for Anthropic, Authorization: Bearer for OpenAI. The gateway forwards it upstream for that one request. It is never written to storage, never logged, and never used for any other caller.
  • Your Revl key (rvl_live_ followed by 32 to 128 URL-safe characters) travels in X-Revl-Key. It only identifies you to the gateway, which stores only the key's SHA-256 digest. It is never forwarded upstream.

Revl AI does not resell, sublicense, pool or share model-provider API access. Every request is authenticated to Anthropic or OpenAI with your own API key, under your own provider account and billing and that provider's terms and usage policies. Revl charges only for the gateway, never for model usage.

API keys only. OAuth and consumer subscription tokens are rejected; only provider API keys are accepted. On the Anthropic route, a request that carries an Authorization header and no x-api-key is rejected with 400 revl_oauth_not_supported.

Quickstart

Keep the SDK you already use. Change the base URL, pass your own provider key as usual, and add one header.

1. Set your keys

You need a Revl key and your own provider API key. The hosted gateway is invite-only, so a hosted Revl key comes with an invitation from the waitlist. If you self-host, you create one yourself. Export the keys so they never appear in code:

Shell
export REVL_API_KEY="rvl_live_your_gateway_key"
export ANTHROPIC_API_KEY="your-anthropic-api-key"   # if you call Anthropic
export OPENAI_API_KEY="your-openai-api-key"         # if you call OpenAI

2. Send a request

Install
pip install anthropic
quickstart.py
import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.proxyrevlvay.com",
    api_key=os.environ["ANTHROPIC_API_KEY"],  # your own Anthropic API key
    default_headers={"X-Revl-Key": os.environ["REVL_API_KEY"]},
)

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)

for block in message.content:
    if block.type == "text":
        print(block.text)
Install
pip install openai
quickstart.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.proxyrevlvay.com/v1",
    api_key=os.environ["OPENAI_API_KEY"],  # your own OpenAI API key
    default_headers={"X-Revl-Key": os.environ["REVL_API_KEY"]},
)

completion = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)

print(completion.choices[0].message.content)
Install
npm install @anthropic-ai/sdk
quickstart.mjs
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://api.proxyrevlvay.com",
  apiKey: process.env.ANTHROPIC_API_KEY, // your own Anthropic API key
  defaultHeaders: { "X-Revl-Key": process.env.REVL_API_KEY },
});

const message = await client.messages.create({
  model: "claude-opus-5-5",
  max_tokens: 16000,
  messages: [{ role: "user", content: "Say hello in one sentence." }],
});

for (const block of message.content) {
  if (block.type === "text") {
    console.log(block.text);
  }
}
Install
npm install openai
quickstart.mjs
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.proxyrevlvay.com/v1",
  apiKey: process.env.OPENAI_API_KEY, // your own OpenAI API key
  defaultHeaders: { "X-Revl-Key": process.env.REVL_API_KEY },
});

const completion = await client.chat.completions.create({
  model: "gpt-4o",
  messages: [{ role: "user", content: "Say hello in one sentence." }],
});

console.log(completion.choices[0].message.content);
Shell
# Anthropic Messages API
curl https://api.proxyrevlvay.com/v1/messages \
  -H "X-Revl-Key: $REVL_API_KEY" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-opus-5-5",
    "max_tokens": 16000,
    "messages": [{"role": "user", "content": "Say hello in one sentence."}]
  }'

# OpenAI Chat Completions API
curl https://api.proxyrevlvay.com/v1/chat/completions \
  -H "X-Revl-Key: $REVL_API_KEY" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}]
  }'
  • Self-hosting? Replace https://api.proxyrevlvay.com with your own Worker URL.
  • The Node.js samples are ES modules that use top-level await. Save them with the .mjs extension.
  • To see what the gateway did, read the X-Revl-* response headers. With cURL, add -i.

Endpoints

Method and pathUpstreamProvider key header (required)SDK base URL
POST /v1/chat/completionshttps://api.openai.com/v1/chat/completionsAuthorization: Bearer <OpenAI API key>OpenAI SDK: <gateway>/v1
POST /v1/messageshttps://api.anthropic.com/v1/messagesx-api-key: <Anthropic API key>Anthropic SDK: <gateway> (the SDK appends /v1/messages)
GET /healthznonenoneReturns {"ok":true,"version":"0.1.0"}

There is no translation layer. OpenAI-format requests go to OpenAI and Anthropic-format requests go to Anthropic. The body is forwarded unchanged unless you turn on masking. On the Anthropic route, the anthropic-version and anthropic-beta request headers are forwarded unchanged.

Request headers reach the provider by allow-list only: content-type, accept, user-agent, x-stainless-*, your provider key header, and openai-organization and openai-project on the OpenAI route. Cookies, every X-Revl-* header and any other client header are dropped, and the query string is not forwarded. A JWT-shaped bearer token on the OpenAI route, or an x-api-key that starts with sk-ant-oat, is rejected with 400 revl_oauth_not_supported.

Request options

Every option is an optional request header. X-Revl-Key and every other X-Revl-* request header are consumed by the gateway and are never forwarded upstream.

HeaderValuesDefaultEffect
X-Revl-Cacheexact, offoffExact-match response cache for this request
X-Revl-Cache-TTLseconds3600Clamped to 60..86400
X-Revl-Maskcomma list of pii, secrets; or offoffMask matches before the request leaves the gateway
X-Revl-Retries0..30Retry upstream on 429, 500, 502, 503, 504 and network errors

An invalid value returns 400 revl_bad_request naming the header.

Example: caching and masking on every request

Python
import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.proxyrevlvay.com",
    api_key=os.environ["ANTHROPIC_API_KEY"],
    default_headers={
        "X-Revl-Key": os.environ["REVL_API_KEY"],
        "X-Revl-Cache": "exact",        # cache identical requests
        "X-Revl-Cache-TTL": "600",      # keep entries for 10 minutes
        "X-Revl-Mask": "pii,secrets",   # mask before the request leaves the gateway
    },
)

To set an option for a single call instead, send the header with that request: extra_headers in the Python SDKs, the headers request option in the JavaScript SDKs.

Response headers

Added to every proxied response.

HeaderValue
X-Revl-Request-Iduuid
X-Revl-Cache-StatusHIT, MISS, or BYPASS when the cache was not used for the request (caching is off, or the request is streamed)
X-Revl-Upstream-Msinteger, 0 on a hit
X-Revl-Maskednumber of distinct values masked
X-Revl-Retriesretries actually used

The upstream status code and body are returned unchanged, including upstream error responses, so each SDK's own error handling keeps working. set-cookie and hop-by-hop headers are dropped from the upstream response.

Caching

Send X-Revl-Cache: exact to cache a response. Matching is exact, not semantic. X-Revl-Cache-TTL sets how long an entry is kept. A hit returns X-Revl-Cache-Status: HIT and X-Revl-Upstream-Ms: 0.

Cache keySHA-256 over the tenant id, the route, the SHA-256 of the provider key, the anthropic-version and anthropic-beta header values, and the canonical JSON of the body that is sent upstream (object keys sorted recursively).
Stored only whenThe request is not streaming, the upstream status is 200, the response body is at most 1 MiB, and the body does not repeat your provider key or gateway key.
What is storedThe status, content type and body of the response, in the KV binding CACHE.
IsolationEntries are scoped to the gateway key (the tenant) and bound to a hash of the provider key, so a cached completion can never be served to a different provider account.

With masking on, the key is computed over the masked body and the stored response is the upstream response before rehydration. On a hit, the response is rehydrated with the current request's mapping. Two prompts that differ only in masked values therefore share one cache entry, which is safe because the provider never saw the values.

Semantic caching: planned design

Roadmap. Not available in v0.1. Exact matching only helps when two requests are identical after canonicalisation. The planned semantic layer reuses a response when a new prompt means the same thing as an earlier one. This is the design we intend to build, published so you can review it before it ships. It may change.

  1. Embed. The final user turn, after masking, is turned into a vector by an embedding model. The system prompt, model name, tools and sampling parameters are not embedded. They are hashed into a strict context key, so a hit can only come from a request with the same instructions and settings.
  2. Search. The vector is looked up in an index partitioned by tenant, context key and the hash of the provider key. Tenants and provider accounts never share a partition.
  3. Decide. The nearest neighbour is used only above a similarity threshold the caller sets per request. Below it, the request goes upstream and the new response is indexed.
  4. Report. A semantic hit gets its own cache status and the similarity score in the response headers, so every reused answer can be audited.

Known limit: semantic reuse can return the answer to a question that only looks similar. It is unsuitable where a small change in wording matters, such as amounts, names or negation. For that reason it will be opt-in per request, with the exact-match cache checked first.

Masking

Send X-Revl-Mask with pii, secrets or pii,secrets. Matches are replaced with placeholders before the request leaves the gateway, and placeholders in the response are replaced with your original values.

Know the limits. Masking is pattern based. It will miss things and is not a compliance control by itself.

Detectors

ValueDetects
piiEmail address; phone number (8 to 15 digits written with a leading +, a bracketed area code, or at least three digit groups; a bare run of digits is not treated as a phone number); payment card (13 to 19 digits that pass the Luhn check); IPv4 address.
secretssk-ant-..., sk-... and sk-proj-... keys; AWS access key ids (AKIA + 16); GitHub tokens (ghp_, gho_, ghu_, ghs_, ghr_, github_pat_); Slack tokens (xox[baprs]-); JSON Web Tokens; PEM private key blocks.

Where it applies

FormatMasked in the requestRehydrated in the response
OpenAImessages[].content (string, or text parts of an array)choices[].message.content
Anthropicsystem (string or text blocks) and messages[].content (string, text blocks, and string or text-block content of tool_result blocks)content[] text blocks

Masking is not applied to tool definitions, tool call arguments, or image and document data. It cannot be combined with streaming in v0.1.

Placeholders and rehydration

Placeholders look like <EMAIL_1>, <PHONE_1>, <CARD_1>, <IP_1> and <SECRET_1>. The same value gets the same placeholder within one request. The mapping exists only in memory for the duration of the request. The X-Revl-Masked response header reports the number of distinct values masked.

Illustrative example with X-Revl-Mask: pii
You sendWrite to jane@example.com and ask whether jane@example.com is still the right address
The provider receivesWrite to <EMAIL_1> and ask whether <EMAIL_1> is still the right address
The provider repliesSubject: Is <EMAIL_1> still the right address?
You receiveSubject: Is jane@example.com still the right address?

Streaming

Set "stream": true in the body as you normally would.

  • The upstream response is passed through unbuffered.
  • Streamed responses are not cached. X-Revl-Cache-Status is BYPASS.
  • Masking with streaming is not supported in v0.1: 400 revl_mask_stream_unsupported.
  • Retries apply only before the first upstream byte.

Rate limits

The gateway limits requests per tenant through the Workers rate limiting binding RATE_LIMITER. The default is 60 requests per 60 seconds, set in wrangler.jsonc. Over the limit you get 429 revl_rate_limited with Retry-After: 60. If the binding is absent, the gateway does not rate limit.

This is separate from your provider's own limits. A provider 429 passes through unchanged, while a gateway 429 has an error type that starts with revl_.

Errors

Every error the gateway itself produces has this shape and carries an X-Revl-Request-Id header. Messages never contain key material.

JSON
{"error": {"type": "revl_<code>", "message": "<human readable>"}}
StatusTypeWhen
400revl_bad_requestAn option header has an invalid value (the message names the header), or the body is not a JSON object
400revl_oauth_not_supportedAnthropic route: an Authorization header and no x-api-key, or an x-api-key that starts with sk-ant-oat. OpenAI route: a JWT-shaped bearer token
400revl_mask_stream_unsupportedMasking is on and the body sets "stream": true
401revl_unauthorizedMissing, unknown or disabled gateway key
401revl_missing_provider_keyMissing provider key on either route
404revl_not_foundAny path that is not listed under Endpoints
405revl_method_not_allowedWrong method on a known path. The response has an Allow header
413revl_payload_too_largeRequest body over 10 MiB
429revl_rate_limitedOver the tenant rate limit. The response has Retry-After: 60
502revl_upstream_unreachableThe upstream could not be reached after all retries
500revl_internal_errorUnexpected failure inside the gateway

Provider errors are not rewritten. The upstream status code and body pass through unchanged.

Self-hosting

The core runs as a Cloudflare Worker with two KV namespaces. You need a Cloudflare account and Node.js. The source is on GitHub. The commands below clone it and deploy the gateway directory.

Shell
# 0. Get the source
git clone https://github.com/trunghuynhnova/revl-ai-gateway.git
cd revl-ai-gateway/gateway

# 1. Install dependencies, sign in to Cloudflare
npm install
npx wrangler login

# 2. Create the two KV namespaces. Each command prints an id.
npx wrangler kv namespace create KEYS
npx wrangler kv namespace create CACHE

# 3. Put each id in the matching "kv_namespaces" entry in wrangler.jsonc, then deploy
npx wrangler deploy

# 4. Create a gateway key. The script prints the key once, and the command that stores its SHA-256 digest.
node scripts/create-key.mjs --id my-team --label "first key"
# Run the "npx wrangler kv key put ..." command it prints, then send the key as X-Revl-Key.

Check the deployment with GET /healthz, which returns {"ok":true,"version":"0.1.0"}. Then follow the quickstart.

SettingKindPurpose
KEYSKV bindingGateway key records, stored under key:<hex> where <hex> is the SHA-256 hex digest of the key
CACHEKV bindingExact-match cache entries
RATE_LIMITERRate limiting bindingPer-tenant rate limit, 60 requests per 60 seconds by default. If absent, the gateway does not rate limit
REQUIRE_GATEWAY_KEYVariableDefaults to "true". When "false" (local development, single-tenant self-hosting), X-Revl-Key is optional and the tenant id is "public"
OPENAI_BASE_URL
ANTHROPIC_BASE_URL
VariablesOverride the upstream, for tests or a regional endpoint. When unset, the upstreams listed under Endpoints are used

Data handling

What the core stores and logs, and what it does not.

DataWhat happens to it
Provider API keyForwarded upstream for that one request. Never written to storage, never logged, never used for any other caller.
Gateway keyOnly its SHA-256 hex digest is stored, in the KV binding KEYS, with a tenant id, a label and a disabled flag.
Request bodyForwarded unchanged unless masking is on. Never logged. With caching on, it is hashed into the SHA-256 cache key; the stored entry holds the response only.
Response bodyNever logged. Stored only as a cache entry, when you turn caching on and the response qualifies.
Masking mappingKept only in memory for the duration of the request.
Request logOne JSON line per request with request_id, tenant, route, model, status, cache, upstream_ms, masked, retries and stream. No bodies, no headers, no key material, no key hashes.

This describes the open-source core. This website collects waitlist email addresses only; see the Privacy Policy.

Available now and roadmap

Roadmap items are not built yet and are in no committed order.

Available in core v0.1

  • BYOK passthrough for the OpenAI Chat Completions API and the Anthropic Messages API
  • Gateway access keys
  • Exact-match response caching
  • PII and secret masking with rehydration (non-streaming)
  • Retries
  • Per-tenant rate limiting
  • Streaming passthrough
  • Metadata-only logging

Roadmap, not built yet

  • Semantic (embedding-based) caching
  • Caching and rehydration for streamed responses
  • Cross-provider fallback
  • Usage dashboard and analytics API
  • Per-key budgets and per-key rate limits
  • SSO and SCIM
  • Custom masking detectors

Found something on this page that does not match what the gateway does? Tell us at founder@proxyrevlvay.com.