Dinesh’sLearning Lab
← The learning library
systems · intermediate · 6 min read
Lesson 11 of 11 in this path ↗

Capstone: an evidence service with replay-safe usage

Compose tenant-scoped retrieval, explicit abstention and a durable-event invariant into one executable failure lab.

Editorial review: · What review means

Stored in this browser only. No account, no sync. Clearing browser data removes your record.

By the end, you should be able to

  • Compose an evidence boundary with idempotent usage
  • Test duplicate and conflicting requests
  • Defend a bounded design using a scored rubric

Bring with you

  • Kimi K3: reconcile the model card before sizing a system

Listen to this article

Browser / device speech · no paid TTS integration. Voice quality depends on your device.

Choose a local device voice to avoid a remote speech service. This site adds no TTS service, account or API calls.

Checking browser speech support…

Pause saves your segment; resume repeats that short segment. Changing voice or speed pauses playback. Stop resets to the beginning. Progress counts finished text segments, not audio time. Leaving or hiding this page stops or pauses speech.

What gets read aloud?

Reads the article body as it appears when you press Listen. Navigation, controls and closed sections are skipped. Expand a section, then Stop and Listen to include it. Code and equations get brief notices; figures use available labels or captions, not their visual details. This narration does not teach omitted mathematics or replace reading examples on the page.

For better sound at no added site cost, try installed English voices, including enhanced voices offered by your device. We cannot guarantee a best voice on every browser. Use Stop or your device’s audio controls if its speech engine misbehaves.

In this article · 6 sections

You now have separate explanations for overload, replay, historical time, ranking and release evidence. The capstone is not to draw all their tools in one diagram. It is to implement the smallest useful boundary and demonstrate what happens when it fails.

Deliverable: a local evidence service function for two synthetic tenants. It retrieves allowed handbook snippets, abstains when none match, and records exactly one accepted-request usage event per tenant-scoped request ID. A repeated ID with a different question is a conflict. There is no network listener, cloud identity provider or model inference. Authentication is a prerequisite represented by a trusted function argument in the fixture.

Choose the product policy before implementation

A new valid request costs one unit, including an abstention: the service still performed retrieval. An exact retry costs zero additional units and returns the original stored packet. This means a retry is tied to the original result, not a newly generated answer. The fixture uses an immutable document snapshot. In production, retained packets require a permission-revocation and retention policy before they can be safely returned again.

The transaction groups the persisted packet with its usage identity. The program does no network call while holding the transaction. A production generator should not occupy a database writer lock while waiting for a remote model; that needs a recoverable job/state-machine design instead.

Complete local implementation

This is a deliberately small reference implementation, not a scaffold with hidden imports. The raw event table is the ledger, so totals are derived by counting accepted unique requests rather than maintaining a second counter.

import json
import re
import sqlite3
 
DOCS = [("a", "a-refund-v1", "Refund requests have a 30 day window."),
        ("b", "b-refund-v1", "Refund requests have a 90 day window.")]
con = sqlite3.connect(":memory:", isolation_level=None)
con.execute("""CREATE TABLE requests(
    tenant TEXT NOT NULL, id TEXT NOT NULL, question TEXT NOT NULL,
    packet TEXT NOT NULL, PRIMARY KEY(tenant, id))""")
 
def words(text):
    return set(re.findall(r"[a-z0-9]+", text.lower()))
 
def serve(tenant, request_id, question, fail=False):
    if not all(isinstance(x, str) and x.strip()
               for x in (tenant, request_id, question)):
        raise ValueError("nonempty strings required")
    con.execute("BEGIN IMMEDIATE")
    try:
        old = con.execute("SELECT question, packet FROM requests WHERE tenant=? AND id=?",
                          (tenant, request_id)).fetchone()
        if old:
            if old[0] != question:
                raise ValueError("conflicting request identity")
            result = json.loads(old[1])
        else:
            hits = [(doc_id, text) for owner, doc_id, text in DOCS
                    if owner == tenant and words(question) & words(text)]
            result = {"status": "evidence" if hits else "abstain",
                      "citations": [doc_id for doc_id, _ in hits],
                      "excerpts": [text for _, text in hits]}
            con.execute("INSERT INTO requests VALUES (?, ?, ?, ?)",
                        (tenant, request_id, question, json.dumps(result)))
            if fail:
                raise RuntimeError("injected pre-commit failure")
        con.execute("COMMIT")
        return result
    except BaseException:
        if con.in_transaction:
            con.execute("ROLLBACK")
        raise
 
try:
    first = serve("a", "r1", "refund")
    assert first["citations"] == ["a-refund-v1"]
    assert serve("a", "r1", "refund") == first
    assert serve("b", "r1", "refund")["citations"] == ["b-refund-v1"]
    assert serve("a", "r2", "quantum")["status"] == "abstain"
    try:
        serve("a", "r1", "different question")
    except ValueError:
        pass
    else:
        raise AssertionError("conflicting ID accepted")
    try:
        serve("a", "r3", "refund", fail=True)
    except RuntimeError:
        pass
    assert con.execute("SELECT count(*) FROM requests WHERE tenant='a'").fetchone()[0] == 2
    assert serve("a", "r3", "refund")["status"] == "evidence"
    counts = con.execute("SELECT tenant, count(*) FROM requests GROUP BY tenant ORDER BY tenant").fetchall()
    assert counts == [("a", 3), ("b", 1)]
    print(counts)
finally:
    con.close()

The injected failure rolls back r3, so its later retry creates one event. The duplicate r1 returns the persisted packet without adding another event. Two tenants may share an ID without sharing evidence. These are observed fixture outcomes, not proof of disk durability or successful client delivery after a network failure.

Required extensions and worked answers

Attempt each extension before reading the answer. Keep the original implementation as a reference and make changes in your own exercise file.

1. Add a capacity decision and an outage policy

Reuse the worksheet assumptions: 240,000 daily requests, six-times peak, 0.24 seconds mean occupied time and 60% utilization imply seven single-request workers in the simplified model. This does not prove the SQLite writer can support that concurrency. Measure transaction duration and contention separately. Choose a bounded queue and reject excess work rather than keeping an unlimited backlog. Document whether a limiter outage denies requests or applies a conservative local quota.

2. Explain what changes when permissions are revoked

The stored packet can contain now-forbidden text. Before returning a retry, reauthorize every referenced source against current permissions or invalidate packets by an authorization/snapshot version. Treat an immutable historical result and present access rights as separate concerns. Do not return a cached answer solely because its request ID matches.

3. Add generation without holding a writer lock across a model call

Persist a request state and enqueue recoverable work, run generation outside the transaction, then persist its result using a guarded state transition. Define stable job identity, duplicate delivery behavior, retries, cancellation, timeouts and the billing event's meaning. The one-function synchronous transaction above does not by itself supply this distributed protocol.

4. Design a release test that cannot pass on missing cases

Enumerate expected IDs for allowed retrieval, tenant isolation, no-answer, exact retry, conflicting retry, rollback and revoked-access cases. Require exact coverage, no privacy violations and explicit human-reviewed support labels for generated answers. Run old and new candidates against the same versioned cases. A syntactically valid citation ID is not a factual-support label.

Assessment rubric

Score each dimension 0, 1 or 2: 0 means missing or contradicted; 1 means explained but not tested; 2 means a runnable positive and negative test plus a stated limitation.

DimensionEvidence required for two points
Identity and permissionCross-tenant and revoked-access cases cannot return unauthorized text
Replay and conflictsExact retry has one effect; changed payload under the same ID is rejected
Transaction failureInjected failure leaves no half-persisted event and retry succeeds
Capacity and windowsUnit-labeled worksheet plus correct fixed/rolling-window counterexample
Evaluation honestyExact case coverage, separate safety gates and uncertainty/sampling limits

The suggested completion target is at least eight out of ten, with full points for identity and replay. This reference program does not earn full production-identity credit: revocation is an extension, not something it secretly implements. Explain that limitation in the submission instead of marking every checkbox complete.

What is deliberately not shipped

There is no authentication service, network endpoint, real LLM, distributed transaction, disk crash test or load test. The outcome is a runnable evidence-and-ledger core plus a defensible plan for its next boundary. That is a useful complete learning artifact, unlike a cloud command list that assumes the hard parts happened elsewhere.

Sources and verification boundary

Primary sources checked on 7 October 2026:

  • SQLite transaction documentation — BEGIN IMMEDIATE acquires a write transaction; SQLite permits one simultaneous writer; explicit COMMIT/ROLLBACK determine transaction boundaries.
  • Apache Kafka 4.1 design: delivery semantics — Output-before-offset can replay; offset-before-output can lose effects; external destinations need cooperation for exactly-once effects.
  • Microsoft: security filter pattern — Security trimming is a query filter, not authentication; trusted application code must apply the permission filter to every query.
  • Google SRE: Handling overload — Request costs vary; resource consumption, quotas, load shedding and retry behavior matter beyond aggregate requests per second.
  • Google SRE: Implementing SLOs — An SLI can be good events divided by eligible events; an error budget is the complement of the objective, over a defined window.

The numerical fixtures and Python programs are teaching examples, not measurements of a deployed service. Python blocks were executed locally; no cloud deployment or model inference was performed.

Previous: Kimi K3: reconcile the model card before sizing a system. Track map.

Pause / Recall / Apply

Can you explain it without the page?

Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.

Stored in this browser only. No account, no sync. Clearing browser data removes your record.