Demo

Simulated gateway dashboard

Change the workload and the gateway options, then watch how exact-match caching, masking, retries and streaming bypass behave. The numbers come from a small model that runs in this page, not from a gateway. The open-source core v0.1 has no dashboard; a usage dashboard is on the roadmap.

Simulated data. Everything on this page is generated in your browser for illustration. It is not live traffic, customer data or a benchmark.

Simulation controls

simulated

Simulation time 0:00

X-Revl-Cache: exact

X-Revl-Mask: off

40%

How often the workload sends an exact copy of an earlier request.

2

Synthetic requests created on each one-second tick.

Key figures

simulated

Window: the charted two minutes of simulated traffic.

Settings in effect: caching exact, masking off, repeat-prompt ratio 40%, 2 requests per second. Every figure follows from these settings and from the illustrative distributions described below. None is a measurement.

Requests
0
Simulated rows in the window.
Cache hit rate
–
Simulated HIT rows divided by all rows. Tracks the repeat-prompt ratio.
Tokens saved
0
Sum of tokens on simulated HIT rows.
Est. cost avoided
$0.00
Tokens saved × $5.00 per million tokens, an illustrative price.
p50 latency
–
Simulated median over all rows, hits and upstream calls together.
p95 latency
–
Simulated 95th percentile over all rows.

Latency: cache hits vs upstream calls

simulated

Median per 5-second interval, in milliseconds, by simulation time. Both latencies are sampled from illustrative distributions, not measured.

  • Cache hit (HIT), solid line
  • Upstream call (MISS or BYPASS), dashed line

View as table
Simulated median latency per 5-second interval
Sim timeCache hit median (ms)Cache hitsUpstream median (ms)Upstream calls

Tokens saved by cache hits

simulated

Tokens per second, averaged over each 5-second interval, by simulation time. Follows the repeat-prompt ratio and request rate set above.

View as table
Simulated tokens saved per 5-second interval
Sim timeTokens saved per secondTokens savedCache hitsRequests

With the chart focused, the left and right arrow keys move a highlight between intervals. The same values are in the table view below the chart.

Request log

simulated

Synthetic rows for tenant demo-tenant, newest first, most recent 50. No prompt or response text appears because the gateway never logs bodies.

Simulated request log for tenant demo-tenant, newest first, most recent 50 rows
Sim timeRequest idRouteModelStatusCache Latency (ms)TokensMaskedRetriesStream

How this simulation works

A few lines of JavaScript in this page play the part of a client workload and of the gateway core v0.1. Nothing is fetched and nothing is stored. Every rate and duration below is an illustrative choice, not a measurement.

Clock and workload
One tick per second creates the chosen number of synthetic requests. With the probability set by the repeat-prompt ratio, a request is an exact copy of one of the 200 most recent distinct request bodies; otherwise it is a new body for one of four example models. 15% of new bodies set "stream": true. On load and on reset the first two minutes are computed at once so the charts start full. The clock stops while paused or while the tab is hidden.
Caching
With X-Revl-Cache: exact, a non-streaming request whose body already has a stored 200 response (default TTL, 3600 seconds) is a HIT. Otherwise it is a MISS, and the response is stored if its status is 200. Streamed requests, and every request while caching is off, are BYPASS.
Latency
A hit takes 12 to 45 ms. An upstream call takes 15 ms plus a sample from a skewed (log-normal) distribution with a 900 ms median, limited to 150 to 8,000 ms.
Retries
Every request sends X-Revl-Retries: 2. Each upstream attempt has a 3% chance of returning 503, which costs 150 to 400 ms and is retried. If the last attempt also fails, the 503 is returned unchanged.
Masking
With X-Revl-Mask: pii,secrets, 35% of bodies contain one to three values a detector would match. The Masked column shows the count and the placeholders, such as <EMAIL_1>, that replace them, and the cache key is computed over the masked body. Streamed requests are sent without masking, because v0.1 rejects that combination with 400 revl_mask_stream_unsupported.
Tokens and cost
Each body has a fixed token count, prompt plus completion, between 300 and 3,000. Tokens saved is the sum over HIT rows. Cost avoided multiplies it by a flat $5.00 per million tokens, which is not any provider's price.
Left out
Gateway keys and rate limiting are not simulated. The core's default limit is 60 requests per 60 seconds per tenant, which this workload would exceed above one request per second. Roadmap items such as semantic caching and cross-provider fallback are not built, so they are not shown.
Log columns
The gateway's own log line holds request_id, tenant, route, model, status, cache, upstream_ms, masked, retries and stream. Token counts, placeholder names and total latency are added here for illustration.

To see the real behaviour, read the documentation or join the waitlist for the hosted gateway.