Skip to content

Integrated platform design

Resilience Gate is a testnet release platform built around a URL-shortening application. Its central use case is to decide whether an exact candidate may move forward using observations that connect application behavior, deployment identity, selected dependency failures, and cleanup. This document joins the existing application’s runtime design with delivery and verification; it is separate from the documentation website’s software design.

The project demonstrates that being healthy before a fault is different from recovering acceptably after one. It combines a real request path with an immutable candidate path, then makes selected observed behavior part of normal staging-to-prod eligibility.

The deployed scope is a private, owned GCP and Radius testnet lab. prod means a production-like profile, not a public production or mainnet target. Dev, staging, prod-like workloads, delivery controllers, and observability share one cluster. Per-environment namespaces and data dependencies organize the system but do not imply cluster, region, or administrative isolation.

System context and major boundaries

Enlarge diagram: System context and major boundaries · Version-controlled diagram source

Requests, artifact delivery, experiment control, and evidence interpretation are separate paths joined by recorded identities. The diagram describes the system; it is not live evidence.

Requirements inferred from source are traceable URL creation/redirect behavior, server-owned payment terms, essential-dependency readiness, optional-cache fallback, digest-based deployment, explicit consequential promotion, bounded experiments, and inspectable evidence. The existing payment, promotion, and failure contracts document the corresponding intended behavior. Retrospective rationale here is source-derived analysis, not an invented historical decision record.

The FastAPI application accepts URL creation and resolves short codes. PostgreSQL is durable storage. Redis accelerates redirects and can fail while the database path remains usable. A dedicated staging signer holds payer private keys for test traffic, while the application asks the configured facilitator to verify and settle authorizations. The signer is a client-side dependency of paid load, not a service the application calls directly to settle payments.

The urls table stores a unique code and destination URL, plus optional settlement identifier, payer, and settled timestamp for paid creation. The settlement identifier is unique so an idempotent facilitator replay cannot be reused to create another URL. These fields belong to the private application audit boundary; public evidence excludes payment identities.

Interface Implemented responsibility
POST /shorten Validate URL, return existing mapping when applicable, enforce configured payment, persist new mapping
GET /{code} Cache lookup, database fallback, temporary redirect or defined error
GET /livez Process liveness, distinct from dependency checks
GET /ready Dependency readiness according to current application policy
GET /metrics Application and request instrumentation
Signer /sign-permit2 Produce a bounded Permit2 authorization/signature without exposing private keys
Facilitator /verify and /settle Validate submitted authorization and settle server-owned requirements

Application handlers, database adapter, and cache adapter are implemented in main.py; facilitator behavior is in payment.py, and the signer service is in signer/main.py. The application design provides the detailed data and request model.

For unpaid creation, the application looks up an existing destination, allocates a code, and persists a new mapping. For paid creation, missing authorization produces a 402 challenge; rejected or invalid authorization produces a generic 402 without the challenge header. The server owns amount, asset, network, recipient, and timeout requirements. A valid submission is verified then settled through the facilitator before the database insert. A successful new creation returns 201; an existing destination can return its mapping without a second charge.

A redirect first uses Redis when available. Ordinary misses and unavailable-cache failures both increment the cache-miss counter and continue to PostgreSQL. Successful database resolution returns 302 and may populate the cache. Redis failure is therefore a degraded acceleration path; PostgreSQL loss affects durable lookup and creation. Liveness remains separate from dependency readiness so dependency failure does not automatically mean the process must restart.

Settlement and persistence are not a distributed transaction. If settlement succeeds and database persistence fails, the payer can have a settlement without a created URL response. Unique settlement storage addresses application reuse, but it does not eliminate that gap. Deliberate reconciliation by validated settlement identifier remains necessary; the platform does not claim automatic resolution of every such failure.

Paid-creation, replay, health, dependency-error, and redirect tests exercise these contracts. A mocked test is evidence of source behavior under its controls, not a substitute for live facilitator or recovery observation.

Trusted-main CI validates, publishes, signs, and verifies an application digest. Kargo discovers it alongside a chart-source revision and records Freight. Dev can promote automatically; staging and prod-like moves are manual. Each Stage renders the selected source and digest into an env/* output commit. Argo CD reconciles that exact desired-state output. Source, image, chart-source, rendered revision, verifier runtime, and run identity remain separate.

Dev verifies readiness. Staging verifies readiness plus the gate: identity/budget checks, Lease acquisition, ready target, paid-success startup marker, serial selected faults, actual Apply evidence, strict returned telemetry, scorecards, and cleanup. Normal staging verification supplies prod eligibility; a separate manual approval is an override rather than upstream proof. Prod-like promotion runs readiness/liveness smoke on its own reconciled target and does not repeat paid load or chaos.

The delivery design details those controller interfaces. The gate design documents state transitions, time windows, and evidence validity. The release journey teaches the complete handoff without reducing every action to CI.

Digest pinning preserves selected bytes; CI signature verification establishes the configured publication identity statement. Kargo and Kubernetes do not independently enforce that signature or consume the CI identity record. Workload Identity and External Secrets avoid long-lived cloud keys in normal paths, but externally provisioned IAM/token correctness remains a live boundary. The identity workflow documents the separate consumers and grants.

Observability provides request/error histograms, dependency state, restart counters, cache behavior, and payment/signer diagnostics. The gate requires valid returned evidence, but selected expressions deliberately use or vector(0). Cache misses alone do not prove an outage. Quiet application errors during signer failure do not establish useful paid traffic because signing can fail before an application request is made.

Passing scorecards are not sufficient for a passing Job: workflow/load outcome and cleanup also matter. Cleanup failure overrides the main verdict, while the source’s single-object absence helper cannot distinguish every read failure from NotFound. Independent recorded post-run readback therefore contributes to the evidence claim rather than an overstated source guarantee.

Design quality Mechanism Practical limit
Traceability Digests and source/render/run identities Artifacts and evidence need retention; histories do not prove future state.
Graceful degradation Redis fallback; clean database-unavailable responses Only selected failure modes are covered.
Process stability Distinct liveness/readiness and restart checks Does not imply every request succeeds during dependency loss.
Controlled experimentation One Pod per selector, serial deadlines, Lease and exact object cleanup Shared cluster and limited fault coverage remain.
Reviewable delivery Kargo output commits; Argo CD reconciliation; manual staging/prod More handoffs and investigation cost than full automation.
Inspectable decisions Per-check scorecards and enclosing analysis records Private detailed bundles reduce public independent inspectability.

The platform does not provide highly available databases/telemetry, multi-region failover, backup restoration, custody guarantees, SLO/on-call assurance, mainnet payment certification, or independent supply-chain admission. These remain explicit known limitations, not hidden assumptions.

The 3 October 2026 report binds the fresh application source 3b708353…, Freight chart-source 0a98c08c…, application digest ab88d89c…, and separate staging/prod renders. Historical October 2 failures and recovery remain attached to their own candidates. Detailed fresh records are privately retained; selected historical scorecards and reviewed screenshots are public.

A later source, controller, configuration, dependency, or credential change needs a fresh relevant verification. The documentation site explains and publishes reviewed records; rebuilding it does not rerun the platform or refresh an old observation’s date. For operations, use the chaos gate runbook, lab lifecycle, and evidence handling, keeping credential-free validation separate from authorized live work.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer