Real-Time Observability
for the AI Agent Era
Apache Doris unifies logs, metrics, traces, and AI agent events on one high-performance analytical engine, so teams troubleshoot faster, control costs, and keep improving AI quality.
Why observabilitymatters forthe agent era.
When teams unify observability across logs, traces, metrics, and AI agent events, five things shift at once:
- Incident detection
- User experience
- Cost at scale
- AI quality
- The link between system behavior and business outcomes
Faster Incident Detection
Logs, traces, and agent events become analyzable execution data, so anomalies and root causes surface before failures spread.
User Experience & SLA
System health maps to real user impact: latency, errors, answer accuracy, task completion, and whether responses are grounded.
Lower Cost at Scale
Hot data stays fast for troubleshooting, aggregates cover trends, and historical data moves to lower-cost storage as telemetry grows.
Continuous AI Improvement
Prompts, responses, RAG context, tool calls, scores, and user feedback show where retrieval, prompting, and task completion can improve.
Business-Aware Operations
Every signal ties back to the customers, tenants, and workflows it affects, and to what AI failures cost in conversion, support load, and revenue.
Already running in production.
Three teams run Apache Doris as the analytical foundation for observability: at scale, on live operational data, across logs, metrics, and events.
MiniMax: PB-Scale Logging on Apache Doris, Off Grafana Loki
After moving off Grafana Loki, every MiniMax business line now logs to an Apache Doris system that serves petabytes with over 99.9% availability and answers queries over billions of log lines in seconds.
- PB-scale log storage with 99.9%+ availability across all business lines
- Keyword and aggregation queries on 1 billion logs return within 2 seconds
- 10 GB/s write throughput with second-level ingestion latency
- 5:1 compression and tiered storage cut storage costs by 70%
NetEase: Elasticsearch and InfluxDB Replaced by Apache Doris
NetEase moved its Eagle monitoring platform off Elasticsearch and its IM time series platform off InfluxDB, and now runs both on Apache Doris with faster queries, less storage, and indexes it can change without rebuilding tables.
- 11× faster queries and 70% lower storage cost than Elasticsearch on monitoring logs
- 67% less storage and half the servers of InfluxDB on time series workloads
- 1 GB/s peak write throughput at up to 1 million TPS
- Inverted indexes added or dropped incrementally, without rewriting tables
Tencent Music: Elasticsearch Replaced, Costs Cut by 80%
Tencent Music Entertainment moved its tag-based search and analytics from Elasticsearch to Apache Doris, where inverted indexes serve full-text search and aggregations in a single SQL query.
- 80% lower overall operational cost than Elasticsearch
- 72% smaller storage footprint (697.7 GB → 195.4 GB on the same dataset)
- 4× faster writes, with ingestion time cut from 10+ hours to under 3 hours
- Alerts down from 20+ per day to single digits per month
What modern observability demandsand how Apache Doris answers.
Five things a modern observability platform has to be good at, and the specific Apache Doris capabilities that meet each one.

