RaboAurora/Trai/Monitoring

Monitoring

Health, SLOs and incidents across every live solution — one operational surface instead of one per platform.

46 healthy1 watch1 degraded
SLO health
96%
2 objectives breaching
Open incidents
4
1 high priority
MTTR
42 min
last 30 days
Error budget
63%
remaining this month

Solution health

Quality, latency and SLO per production solution.

SolutionOwnerQualityLatencySLOStatus
Customer 360 Data ProductRetail Data Products97%182 ms99.9%Healthy
CRD Customer Risk FeaturesRisk Analytics96%96 ms99.8%Healthy
Mortgage Risk Model v2.3Risk Analytics94%182 ms99.7%Watch
Onboarding support agentRetail Digital93%2.4 s99.5%Healthy
Credit Exposure MonitoringRisk Analytics94%1.1 s99.9%Healthy
Segment Activation ServiceMarketing CDP90%620 ms99.2%Degraded

Service level objectives

Declared in the data contract, measured here.

ObjectiveTargetActualStatus
Availability99.9%99.94%meeting
Freshness p9515 min7 minmeeting
API latency p95250 ms182 msmeeting
Data quality95%96.7%meeting
Agent resolution75%81%meeting
Activation latency p95500 ms620 msbreaching

Incidents

Freshness warning
Customer 360 · Source feed delayed 11 minutes · 18m
Investigating
Activation latency breach
Segment Activation · p95 above 500 ms SLO · 52m
Mitigating
Quality rule drift
Customer 360 · Null ratio above baseline · 3h
Review
Cost anomaly
Onboarding agent · Compute usage +12% vs plan · 1d
Watch

Drift & quality watch

Signals that precede an incident.

Feature driftwithin tolerance
Prediction driftwithin tolerance
Null ratioabove baseline
Volume anomalywithin tolerance
Cost per run+12% vs plan