How BYOK works
BYOK means bring your own key. The gateway holds no model-provider credentials. Every request carries two keys, and they do different jobs.
Your app
Sends both keys as request headers.
Revl Gateway
Checks the Revl key, then drops it. Never stores or logs the provider key.
Provider
Anthropic or OpenAI, under your account, your billing and their terms.
- Your provider API key travels in that provider's standard header:
x-api-keyfor Anthropic,Authorization: Bearerfor OpenAI. The gateway forwards it upstream for that one request. It is never written to storage, never logged, and never used for any other caller. - Your Revl key (
rvl_live_followed by 32 to 128 URL-safe characters) travels inX-Revl-Key. It only identifies you to the gateway, which stores only the key's SHA-256 digest. It is never forwarded upstream.
Revl AI does not resell, sublicense, pool or share model-provider API access. Every request is authenticated to Anthropic or OpenAI with your own API key, under your own provider account and billing and that provider's terms and usage policies. Revl charges only for the gateway, never for model usage.
API keys only. OAuth and consumer subscription tokens are rejected; only provider API keys are accepted. On the Anthropic route, a request that carries an Authorization header and no x-api-key is rejected with 400 revl_oauth_not_supported.
Quickstart
Keep the SDK you already use. Change the base URL, pass your own provider key as usual, and add one header.
1. Set your keys
You need a Revl key and your own provider API key. The hosted gateway is invite-only, so a hosted Revl key comes with an invitation from the waitlist. If you self-host, you create one yourself. Export the keys so they never appear in code:
export REVL_API_KEY="rvl_live_your_gateway_key"
export ANTHROPIC_API_KEY="your-anthropic-api-key" # if you call Anthropic
export OPENAI_API_KEY="your-openai-api-key" # if you call OpenAI
2. Send a request
pip install anthropic
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.proxyrevlvay.com",
api_key=os.environ["ANTHROPIC_API_KEY"], # your own Anthropic API key
default_headers={"X-Revl-Key": os.environ["REVL_API_KEY"]},
)
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
for block in message.content:
if block.type == "text":
print(block.text)
pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.proxyrevlvay.com/v1",
api_key=os.environ["OPENAI_API_KEY"], # your own OpenAI API key
default_headers={"X-Revl-Key": os.environ["REVL_API_KEY"]},
)
completion = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(completion.choices[0].message.content)
npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
baseURL: "https://api.proxyrevlvay.com",
apiKey: process.env.ANTHROPIC_API_KEY, // your own Anthropic API key
defaultHeaders: { "X-Revl-Key": process.env.REVL_API_KEY },
});
const message = await client.messages.create({
model: "claude-opus-5-5",
max_tokens: 16000,
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
for (const block of message.content) {
if (block.type === "text") {
console.log(block.text);
}
}
npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.proxyrevlvay.com/v1",
apiKey: process.env.OPENAI_API_KEY, // your own OpenAI API key
defaultHeaders: { "X-Revl-Key": process.env.REVL_API_KEY },
});
const completion = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(completion.choices[0].message.content);
# Anthropic Messages API
curl https://api.proxyrevlvay.com/v1/messages \
-H "X-Revl-Key: $REVL_API_KEY" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 16000,
"messages": [{"role": "user", "content": "Say hello in one sentence."}]
}'
# OpenAI Chat Completions API
curl https://api.proxyrevlvay.com/v1/chat/completions \
-H "X-Revl-Key: $REVL_API_KEY" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Say hello in one sentence."}]
}'
- Self-hosting? Replace
https://api.proxyrevlvay.comwith your own Worker URL. - The Node.js samples are ES modules that use top-level
await. Save them with the.mjsextension. - To see what the gateway did, read the
X-Revl-*response headers. With cURL, add-i.
Endpoints
| Method and path | Upstream | Provider key header (required) | SDK base URL |
|---|---|---|---|
POST /v1/chat/completions | https://api.openai.com/v1/chat/completions | Authorization: Bearer <OpenAI API key> | OpenAI SDK: <gateway>/v1 |
POST /v1/messages | https://api.anthropic.com/v1/messages | x-api-key: <Anthropic API key> | Anthropic SDK: <gateway> (the SDK appends /v1/messages) |
GET /healthz | none | none | Returns {"ok":true,"version":"0.1.0"} |
There is no translation layer. OpenAI-format requests go to OpenAI and Anthropic-format requests go to Anthropic. The body is forwarded unchanged unless you turn on masking. On the Anthropic route, the anthropic-version and anthropic-beta request headers are forwarded unchanged.
Request headers reach the provider by allow-list only: content-type, accept, user-agent, x-stainless-*, your provider key header, and openai-organization and openai-project on the OpenAI route. Cookies, every X-Revl-* header and any other client header are dropped, and the query string is not forwarded. A JWT-shaped bearer token on the OpenAI route, or an x-api-key that starts with sk-ant-oat, is rejected with 400 revl_oauth_not_supported.
Request options
Every option is an optional request header. X-Revl-Key and every other X-Revl-* request header are consumed by the gateway and are never forwarded upstream.
| Header | Values | Default | Effect |
|---|---|---|---|
X-Revl-Cache | exact, off | off | Exact-match response cache for this request |
X-Revl-Cache-TTL | seconds | 3600 | Clamped to 60..86400 |
X-Revl-Mask | comma list of pii, secrets; or off | off | Mask matches before the request leaves the gateway |
X-Revl-Retries | 0..3 | 0 | Retry upstream on 429, 500, 502, 503, 504 and network errors |
An invalid value returns 400 revl_bad_request naming the header.
Example: caching and masking on every request
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.proxyrevlvay.com",
api_key=os.environ["ANTHROPIC_API_KEY"],
default_headers={
"X-Revl-Key": os.environ["REVL_API_KEY"],
"X-Revl-Cache": "exact", # cache identical requests
"X-Revl-Cache-TTL": "600", # keep entries for 10 minutes
"X-Revl-Mask": "pii,secrets", # mask before the request leaves the gateway
},
)
To set an option for a single call instead, send the header with that request: extra_headers in the Python SDKs, the headers request option in the JavaScript SDKs.
Response headers
Added to every proxied response.
| Header | Value |
|---|---|
X-Revl-Request-Id | uuid |
X-Revl-Cache-Status | HIT, MISS, or BYPASS when the cache was not used for the request (caching is off, or the request is streamed) |
X-Revl-Upstream-Ms | integer, 0 on a hit |
X-Revl-Masked | number of distinct values masked |
X-Revl-Retries | retries actually used |
The upstream status code and body are returned unchanged, including upstream error responses, so each SDK's own error handling keeps working. set-cookie and hop-by-hop headers are dropped from the upstream response.
Caching
Send X-Revl-Cache: exact to cache a response. Matching is exact, not semantic. X-Revl-Cache-TTL sets how long an entry is kept. A hit returns X-Revl-Cache-Status: HIT and X-Revl-Upstream-Ms: 0.
| Cache key | SHA-256 over the tenant id, the route, the SHA-256 of the provider key, the anthropic-version and anthropic-beta header values, and the canonical JSON of the body that is sent upstream (object keys sorted recursively). |
|---|---|
| Stored only when | The request is not streaming, the upstream status is 200, the response body is at most 1 MiB, and the body does not repeat your provider key or gateway key. |
| What is stored | The status, content type and body of the response, in the KV binding CACHE. |
| Isolation | Entries are scoped to the gateway key (the tenant) and bound to a hash of the provider key, so a cached completion can never be served to a different provider account. |
With masking on, the key is computed over the masked body and the stored response is the upstream response before rehydration. On a hit, the response is rehydrated with the current request's mapping. Two prompts that differ only in masked values therefore share one cache entry, which is safe because the provider never saw the values.
Semantic caching: planned design
Roadmap. Not available in v0.1. Exact matching only helps when two requests are identical after canonicalisation. The planned semantic layer reuses a response when a new prompt means the same thing as an earlier one. This is the design we intend to build, published so you can review it before it ships. It may change.
- Embed. The final user turn, after masking, is turned into a vector by an embedding model. The system prompt, model name, tools and sampling parameters are not embedded. They are hashed into a strict context key, so a hit can only come from a request with the same instructions and settings.
- Search. The vector is looked up in an index partitioned by tenant, context key and the hash of the provider key. Tenants and provider accounts never share a partition.
- Decide. The nearest neighbour is used only above a similarity threshold the caller sets per request. Below it, the request goes upstream and the new response is indexed.
- Report. A semantic hit gets its own cache status and the similarity score in the response headers, so every reused answer can be audited.
Known limit: semantic reuse can return the answer to a question that only looks similar. It is unsuitable where a small change in wording matters, such as amounts, names or negation. For that reason it will be opt-in per request, with the exact-match cache checked first.
Masking
Send X-Revl-Mask with pii, secrets or pii,secrets. Matches are replaced with placeholders before the request leaves the gateway, and placeholders in the response are replaced with your original values.
Know the limits. Masking is pattern based. It will miss things and is not a compliance control by itself.
Detectors
| Value | Detects |
|---|---|
pii | Email address; phone number (8 to 15 digits written with a leading +, a bracketed area code, or at least three digit groups; a bare run of digits is not treated as a phone number); payment card (13 to 19 digits that pass the Luhn check); IPv4 address. |
secrets | sk-ant-..., sk-... and sk-proj-... keys; AWS access key ids (AKIA + 16); GitHub tokens (ghp_, gho_, ghu_, ghs_, ghr_, github_pat_); Slack tokens (xox[baprs]-); JSON Web Tokens; PEM private key blocks. |
Where it applies
| Format | Masked in the request | Rehydrated in the response |
|---|---|---|
| OpenAI | messages[].content (string, or text parts of an array) | choices[].message.content |
| Anthropic | system (string or text blocks) and messages[].content (string, text blocks, and string or text-block content of tool_result blocks) | content[] text blocks |
Masking is not applied to tool definitions, tool call arguments, or image and document data. It cannot be combined with streaming in v0.1.
Placeholders and rehydration
Placeholders look like <EMAIL_1>, <PHONE_1>, <CARD_1>, <IP_1> and <SECRET_1>. The same value gets the same placeholder within one request. The mapping exists only in memory for the duration of the request. The X-Revl-Masked response header reports the number of distinct values masked.
Illustrative example with X-Revl-Mask: pii | |
|---|---|
| You send | Write to jane@example.com and ask whether jane@example.com is still the right address |
| The provider receives | Write to <EMAIL_1> and ask whether <EMAIL_1> is still the right address |
| The provider replies | Subject: Is <EMAIL_1> still the right address? |
| You receive | Subject: Is jane@example.com still the right address? |
Streaming
Set "stream": true in the body as you normally would.
- The upstream response is passed through unbuffered.
- Streamed responses are not cached.
X-Revl-Cache-StatusisBYPASS. - Masking with streaming is not supported in v0.1:
400 revl_mask_stream_unsupported. - Retries apply only before the first upstream byte.
Rate limits
The gateway limits requests per tenant through the Workers rate limiting binding RATE_LIMITER. The default is 60 requests per 60 seconds, set in wrangler.jsonc. Over the limit you get 429 revl_rate_limited with Retry-After: 60. If the binding is absent, the gateway does not rate limit.
This is separate from your provider's own limits. A provider 429 passes through unchanged, while a gateway 429 has an error type that starts with revl_.
Errors
Every error the gateway itself produces has this shape and carries an X-Revl-Request-Id header. Messages never contain key material.
{"error": {"type": "revl_<code>", "message": "<human readable>"}}
| Status | Type | When |
|---|---|---|
| 400 | revl_bad_request | An option header has an invalid value (the message names the header), or the body is not a JSON object |
| 400 | revl_oauth_not_supported | Anthropic route: an Authorization header and no x-api-key, or an x-api-key that starts with sk-ant-oat. OpenAI route: a JWT-shaped bearer token |
| 400 | revl_mask_stream_unsupported | Masking is on and the body sets "stream": true |
| 401 | revl_unauthorized | Missing, unknown or disabled gateway key |
| 401 | revl_missing_provider_key | Missing provider key on either route |
| 404 | revl_not_found | Any path that is not listed under Endpoints |
| 405 | revl_method_not_allowed | Wrong method on a known path. The response has an Allow header |
| 413 | revl_payload_too_large | Request body over 10 MiB |
| 429 | revl_rate_limited | Over the tenant rate limit. The response has Retry-After: 60 |
| 502 | revl_upstream_unreachable | The upstream could not be reached after all retries |
| 500 | revl_internal_error | Unexpected failure inside the gateway |
Provider errors are not rewritten. The upstream status code and body pass through unchanged.
Self-hosting
The core runs as a Cloudflare Worker with two KV namespaces. You need a Cloudflare account and Node.js. The source is on GitHub. The commands below clone it and deploy the gateway directory.
# 0. Get the source
git clone https://github.com/trunghuynhnova/revl-ai-gateway.git
cd revl-ai-gateway/gateway
# 1. Install dependencies, sign in to Cloudflare
npm install
npx wrangler login
# 2. Create the two KV namespaces. Each command prints an id.
npx wrangler kv namespace create KEYS
npx wrangler kv namespace create CACHE
# 3. Put each id in the matching "kv_namespaces" entry in wrangler.jsonc, then deploy
npx wrangler deploy
# 4. Create a gateway key. The script prints the key once, and the command that stores its SHA-256 digest.
node scripts/create-key.mjs --id my-team --label "first key"
# Run the "npx wrangler kv key put ..." command it prints, then send the key as X-Revl-Key.
Check the deployment with GET /healthz, which returns {"ok":true,"version":"0.1.0"}. Then follow the quickstart.
| Setting | Kind | Purpose |
|---|---|---|
KEYS | KV binding | Gateway key records, stored under key:<hex> where <hex> is the SHA-256 hex digest of the key |
CACHE | KV binding | Exact-match cache entries |
RATE_LIMITER | Rate limiting binding | Per-tenant rate limit, 60 requests per 60 seconds by default. If absent, the gateway does not rate limit |
REQUIRE_GATEWAY_KEY | Variable | Defaults to "true". When "false" (local development, single-tenant self-hosting), X-Revl-Key is optional and the tenant id is "public" |
OPENAI_BASE_URLANTHROPIC_BASE_URL | Variables | Override the upstream, for tests or a regional endpoint. When unset, the upstreams listed under Endpoints are used |
Data handling
What the core stores and logs, and what it does not.
| Data | What happens to it |
|---|---|
| Provider API key | Forwarded upstream for that one request. Never written to storage, never logged, never used for any other caller. |
| Gateway key | Only its SHA-256 hex digest is stored, in the KV binding KEYS, with a tenant id, a label and a disabled flag. |
| Request body | Forwarded unchanged unless masking is on. Never logged. With caching on, it is hashed into the SHA-256 cache key; the stored entry holds the response only. |
| Response body | Never logged. Stored only as a cache entry, when you turn caching on and the response qualifies. |
| Masking mapping | Kept only in memory for the duration of the request. |
| Request log | One JSON line per request with request_id, tenant, route, model, status, cache, upstream_ms, masked, retries and stream. No bodies, no headers, no key material, no key hashes. |
This describes the open-source core. This website collects waitlist email addresses only; see the Privacy Policy.
Available now and roadmap
Roadmap items are not built yet and are in no committed order.
Available in core v0.1
- BYOK passthrough for the OpenAI Chat Completions API and the Anthropic Messages API
- Gateway access keys
- Exact-match response caching
- PII and secret masking with rehydration (non-streaming)
- Retries
- Per-tenant rate limiting
- Streaming passthrough
- Metadata-only logging
Roadmap, not built yet
- Semantic (embedding-based) caching
- Caching and rehydration for streamed responses
- Cross-provider fallback
- Usage dashboard and analytics API
- Per-key budgets and per-key rate limits
- SSO and SCIM
- Custom masking detectors
Found something on this page that does not match what the gateway does? Tell us at founder@proxyrevlvay.com.