Metrics and scoring rules
The implemented rules live in the chaos scorer and baseline scorer. This reference preserves their exact query semantics. The bounds are project verification rules, not an availability SLO or a public production guarantee.
Metric sources and meanings
Section titled “Metric sources and meanings”The ServiceMonitor selects the URL-shortener Service in each reviewed application namespace, scrapes /metrics every 15 seconds with a 10-second timeout, and attaches a namespace label. Application HTTP instrumentation excludes /livez, /ready, and /metrics as configured in application source. Kubernetes readiness and restart signals come from the observability stack’s Kubernetes state metrics.
| Metric | Meaning relevant to review |
|---|---|
http_requests_total |
Instrumented HTTP request count by route/status/method labels; it does not alone identify a paid success. |
http_request_duration_seconds_bucket |
HTTP duration histogram in seconds; scorer uses route-template labels. |
url_shortener_cache_hits_total |
Redirects resolved through a cached value. |
url_shortener_cache_misses_total |
Ordinary cache misses or unavailable-cache fallbacks; it is not an outage-only counter. |
url_shortener_urls_created_total |
Successfully persisted new short URLs. |
url_shortener_payment_outcomes_total{outcome=…} |
Coarse payment decision outcomes, not raw facilitator messages. |
url_shortener_payment_replays_total |
Settlement identifiers rejected as already used by the application. |
url_shortener_dependency_up{dependency=…} |
Dependency reachability at the most recent health check; not the Kubernetes target-readiness expression used for chaos. |
url_shortener_facilitator_calls_total{endpoint=…,outcome=…} |
Facilitator verify/settle call results. |
url_shortener_facilitator_request_duration_seconds |
Per-facilitator-call duration histogram. |
kube_pod_status_ready{condition="true"} |
Kubernetes Pod readiness state for the named target. |
kube_pod_container_status_restarts_total |
Kubernetes restart count for the application container. |
Custom metric definitions are in main.py and payment.py; metric tests and payment-metric tests cover their behavior. Payment dashboard display and gate assertions are different consumers of telemetry.
Windows, freshness, and query kinds
Section titled “Windows, freshness, and query kinds”Let T be the earliest successful Apply event on an experiment’s exact PodChaos child. The runner invokes each scorer with duration 60 seconds; each definition uses settle 60 seconds. Therefore:
| Item | Current meaning |
|---|---|
| Range start | T − 20 seconds |
| Range end / instant evaluation time | T + 120 seconds |
Counter lookback W |
120s, fault plus scored settling; no lead included |
| Range step | 15 seconds |
| Prometheus request timeout | 15 seconds |
| Minimum samples | 3 per returned range series |
| Allowed range start/end distance | First sample no later than start + 60s; last sample no earlier than end − 60s |
| Maximum adjacent range gap | 75 seconds |
| Instant freshness | Absolute timestamp distance from evaluation time ≤ 60 seconds |
The workflow’s PostgreSQL recovery pause is 120 seconds; its scorer still uses 60 seconds of settling. Redis and signer workflow recovery pauses are 60 seconds. A full workflow screenshot interval is not interchangeable with these query windows.
Instant results must contain exactly one sample after the Prometheus expression’s aggregation. Range aggregations operate over all validated returned samples: min, max, and last below mean the Python scorer aggregation, following the PromQL expression. Equality uses absolute tolerance 1e-9; strict comparisons remain strict.
Empty results, wrong response/result shapes, failed transport, non-finite values, unordered timestamps, stale data, or insufficient coverage fail the affected check. The deliberate query fallback or vector(0) can produce a valid returned zero when its left side is absent; the parser does not fail that synthetic result merely because the underlying counter is absent.
PostgreSQL checks
Section titled “PostgreSQL checks”The table’s queries show the current staging namespace and W=120s. A namespace CLI override changes the label value, but exact target names in the PostgreSQL and Redis definitions remain staging names.
| Check ID | Query kind / scorer aggregation | Pass rule and unit |
|---|---|---|
meaningful-traffic |
Instant / value | > 50 requests |
app-restarts |
Instant / value | < 0.5 restarts |
postgres-outage-observed |
Range / min | == 0 state |
postgres-recovered |
Range / last | == 1 state |
postgres-clean-degradation |
Instant / value | < 0.5 responses |
# meaningful-trafficsum(increase(http_requests_total{namespace="url-shortener-staging", handler!~"/(livez|ready|metrics)"}[120s]))# app-restartsmax(increase(kube_pod_container_status_restarts_total{namespace="url-shortener-staging", container="url-shortener"}[120s])) or vector(0)# postgres-outage-observed and postgres-recoveredmin(kube_pod_status_ready{namespace="url-shortener-staging", condition="true", pod="url-shortener-staging-postgresql-0"})# postgres-clean-degradationsum(increase(http_requests_total{namespace="url-shortener-staging", handler="/shorten", status="500"}[120s])) or vector(0)Clean degradation checks the specific /shorten status 500 symptom. It does not require all requests to succeed and does not ban every 5xx status; expected database unavailability can produce 503. Traffic totals are non-probe requests, not a dedicated count of paid creations. PostgreSQL scorer tests cover landed fault and recovery requirements.
Redis checks
Section titled “Redis checks”| Check ID | Query kind / scorer aggregation | Pass rule and unit |
|---|---|---|
meaningful-traffic |
Instant / value | > 50 requests |
redis-fallback-observed |
Instant / value | > 0 fallbacks |
redirect-latency |
Range / max | < 1.2 seconds |
app-restarts |
Instant / value | < 0.5 restarts |
redis-outage-observed |
Range / min | == 0 state |
redis-recovered |
Range / last | == 1 state |
# meaningful-trafficsum(increase(http_requests_total{namespace="url-shortener-staging", handler!~"/(livez|ready|metrics)"}[120s]))# redis-fallback-observedsum(increase(url_shortener_cache_misses_total{namespace="url-shortener-staging"}[120s]))# redirect-latency: maximum of returned one-minute p95 valueshistogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{namespace="url-shortener-staging", handler="/{code}"}[1m])) by (le))# app-restartsmax(increase(kube_pod_container_status_restarts_total{namespace="url-shortener-staging", container="url-shortener"}[120s])) or vector(0)# redis-outage-observed and redis-recoveredmin(kube_pod_status_ready{namespace="url-shortener-staging", condition="true", pod="url-shortener-staging-redis-master-0"})The maximum range p95 is not the p95 of all requests pooled over the 140-second range. Each returned point uses a one-minute rate. The positive cache-miss count must be interpreted alongside observed Redis outage; ordinary misses also increment it. Redis scorer tests check slow fallback, missing fallback evidence, route-label preservation, and exact target selection.
Signer checks
Section titled “Signer checks”| Check ID | Query kind / scorer aggregation | Pass rule and unit |
|---|---|---|
signer-outage-observed |
Range / min | == 0 state |
signer-recovered |
Range / last | == 1 state |
app-restarts |
Instant / value | < 0.5 restarts |
app-5xx |
Instant / value | < 0.5 responses |
payment-replays |
Instant / value | < 0.5 replays |
# signer-outage-observed and signer-recoveredmin(kube_pod_status_ready{namespace="url-shortener-staging", condition="true", pod=~"radius-signer-[a-z0-9]+-[a-z0-9]+"})# app-restartsmax(increase(kube_pod_container_status_restarts_total{namespace="url-shortener-staging", container="url-shortener"}[120s])) or vector(0)# app-5xxsum(increase(http_requests_total{namespace="url-shortener-staging", handler="/shorten", status=~"5.."}[120s])) or vector(0)# payment-replayssum(increase(url_shortener_payment_replays_total{namespace="url-shortener-staging"}[120s])) or vector(0)There is no signer meaningful-traffic check. The load client returns early when it cannot obtain a signature, so low application errors cannot establish continuous paid useful work. Signer scorer tests check dependency recovery and collateral application symptoms.
Separate dev baseline rules
Section titled “Separate dev baseline rules”The baseline scorer uses explicit start/end times rather than a manufactured injection time. Its observation duration must be an integer from 60 to 150 seconds. W below is that exact duration. Instant queries evaluate at the supplied end; range queries span the supplied window and share the coverage validation above.
| Check ID | Query / evaluation | Pass rule |
|---|---|---|
root-get-traffic |
sum(increase(http_requests_total{namespace="N", handler="/", method="GET"}[Ws])), instant/value |
> 20 requests |
application-5xx |
sum(increase(http_requests_total{namespace="N", status=~"5.."}[Ws])) or vector(0), instant/value |
< 0.5 responses |
root-get-p95-latency |
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{namespace="N", handler="/", method="GET"}[1m])) by (le)), instant/value |
< 1.0 seconds |
postgres-ready |
min(kube_pod_status_ready{namespace="N", condition="true", pod=~"url-shortener-[a-z0-9-]+-postgresql-0"}), range/min |
== 1 state |
redis-ready |
min(kube_pod_status_ready{namespace="N", condition="true", pod=~"url-shortener-[a-z0-9-]+-redis-master-0"}), range/min |
== 1 state |
Baseline p95 is one fresh instant value covering the final minute; unlike Redis chaos latency, it is not a maximum across range points. The baseline produces resilience-gate.baseline-scorecard/v1, distinct from the three chaos experiment cards. Baseline tests cover fresh evidence, invalid windows/identities, and sanitized failure output.
Metadata and enclosing outcomes
Section titled “Metadata and enclosing outcomes”The chaos CLI appends required release-identity validation. It needs a source/Freight revision, lowercase sha256 digest, and a bounded safe run ID. Its card verdict passes only when every included check passes. Query failures retain observed: null, an evidence error, and a reason; they are not represented as a guessed zero.
The scorer does not decide cleanup or Kargo promotion. Those belong to gate orchestration and score-to-promotion interpretation. Read query tests, no-data tests, and scorecard tests when updating this reference against future source.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer