Demo
Simulated gateway dashboard
Change the workload and the gateway options, then watch how exact-match caching, masking, retries and streaming bypass behave. The numbers come from a small model that runs in this page, not from a gateway. The open-source core v0.1 has no dashboard; a usage dashboard is on the roadmap.
Simulated data. Everything on this page is generated in your browser for illustration. It is not live traffic, customer data or a benchmark.
Simulation controls
simulatedSimulation time 0:00
X-Revl-Cache: exact
X-Revl-Mask: off
How often the workload sends an exact copy of an earlier request.
Synthetic requests created on each one-second tick.
Key figures
simulatedWindow: the charted two minutes of simulated traffic.
Settings in effect: caching exact, masking off, repeat-prompt ratio 40%, 2 requests per second. Every figure follows from these settings and from the illustrative distributions described below. None is a measurement.
- Requests
- 0
- Simulated rows in the window.
- Cache hit rate
- –
- Simulated HIT rows divided by all rows. Tracks the repeat-prompt ratio.
- Tokens saved
- 0
- Sum of tokens on simulated HIT rows.
- Est. cost avoided
- $0.00
- Tokens saved × $5.00 per million tokens, an illustrative price.
- p50 latency
- –
- Simulated median over all rows, hits and upstream calls together.
- p95 latency
- –
- Simulated 95th percentile over all rows.
Latency: cache hits vs upstream calls
simulatedMedian per 5-second interval, in milliseconds, by simulation time. Both latencies are sampled from illustrative distributions, not measured.
- Cache hit (HIT), solid line
- Upstream call (MISS or BYPASS), dashed line
View as table
| Sim time | Cache hit median (ms) | Cache hits | Upstream median (ms) | Upstream calls |
|---|
Tokens saved by cache hits
simulatedTokens per second, averaged over each 5-second interval, by simulation time. Follows the repeat-prompt ratio and request rate set above.
View as table
| Sim time | Tokens saved per second | Tokens saved | Cache hits | Requests |
|---|
With the chart focused, the left and right arrow keys move a highlight between intervals. The same values are in the table view below the chart.
Request log
simulatedSynthetic rows for tenant demo-tenant, newest first, most recent 50. No prompt or response text appears because the gateway never logs bodies.
| Sim time | Request id | Route | Model | Status | Cache | Latency (ms) | Tokens | Masked | Retries | Stream |
|---|
How this simulation works
A few lines of JavaScript in this page play the part of a client workload and of the gateway core v0.1. Nothing is fetched and nothing is stored. Every rate and duration below is an illustrative choice, not a measurement.
- Clock and workload
- One tick per second creates the chosen number of synthetic requests. With the probability set by the repeat-prompt ratio, a request is an exact copy of one of the 200 most recent distinct request bodies; otherwise it is a new body for one of four example models. 15% of new bodies set
"stream": true. On load and on reset the first two minutes are computed at once so the charts start full. The clock stops while paused or while the tab is hidden. - Caching
- With
X-Revl-Cache: exact, a non-streaming request whose body already has a stored 200 response (default TTL, 3600 seconds) is aHIT. Otherwise it is aMISS, and the response is stored if its status is 200. Streamed requests, and every request while caching is off, areBYPASS. - Latency
- A hit takes 12 to 45 ms. An upstream call takes 15 ms plus a sample from a skewed (log-normal) distribution with a 900 ms median, limited to 150 to 8,000 ms.
- Retries
- Every request sends
X-Revl-Retries: 2. Each upstream attempt has a 3% chance of returning 503, which costs 150 to 400 ms and is retried. If the last attempt also fails, the 503 is returned unchanged. - Masking
- With
X-Revl-Mask: pii,secrets, 35% of bodies contain one to three values a detector would match. The Masked column shows the count and the placeholders, such as<EMAIL_1>, that replace them, and the cache key is computed over the masked body. Streamed requests are sent without masking, because v0.1 rejects that combination with400 revl_mask_stream_unsupported. - Tokens and cost
- Each body has a fixed token count, prompt plus completion, between 300 and 3,000. Tokens saved is the sum over
HITrows. Cost avoided multiplies it by a flat $5.00 per million tokens, which is not any provider's price. - Left out
- Gateway keys and rate limiting are not simulated. The core's default limit is 60 requests per 60 seconds per tenant, which this workload would exceed above one request per second. Roadmap items such as semantic caching and cross-provider fallback are not built, so they are not shown.
- Log columns
- The gateway's own log line holds
request_id,tenant,route,model,status,cache,upstream_ms,masked,retriesandstream. Token counts, placeholder names and total latency are added here for illustration.
To see the real behaviour, read the documentation or join the waitlist for the hosted gateway.