Full-Stack Infrastructure Observability

DRSHYM

To see everything.

Drshym is Ureka AI's observability platform for the complete AI infrastructure stack, from tensor cores and kernels to clusters, racks, power systems, cooling plants, and datacenter operations.

Global AI Infrastructure

us-west / 4,096 GPUs / 512 MW campus
GPU Utilization87.4%+8.1%
HBM Bandwidth3.18 TB/s92% peak
NVLink Saturation71.8%healthy
Tokens / sec2.84M+12.4%
Rack Power118 kWwithin cap
Campus PUE1.17optimized

End-to-End Latency Correlation

Rack Thermal Map

Live Signal Stream

21:44:08.114 gpu.122 sm_active=91.2%
21:44:08.119 hbm.stack2 queue_depth=27
21:44:08.126 nvlink.7 retry_rate=0.018%
21:44:08.131 pod.38 ttft_p99=118ms
21:44:08.142 rack.R17 inlet_temp=24.8C
21:44:08.151 cdu.04 delta_p=2.14bar
21:44:08.164 ups.B output_load=76.3%
21:44:08.181 chiller.02 cop=6.41

Infrastructure Health

Accelerators99.7%
Fabric99.9%
Storage99.5%
Cooling99.8%
Power100%

Active Anomalies

HBM queue buildup3
Thermal drift2
Link degradation1
Scheduler skew4

Illustrative product view — sample data, not a live customer environment.

One observability plane

From transistor behavior to datacenter economics.

Drshym correlates hardware counters, runtime signals, model behavior, cluster state, physical infrastructure, and business outcomes in one system of record.

01 / SILICON

Hardware & Accelerators

See how the silicon behaves under real production workloads.

  • SM, tensor core and SIMD activity
  • HBM bandwidth, queues and conflicts
  • PCIe, NVLink, CXL and NIC telemetry
  • ECC, throttling, clocks and voltage
  • Chip, board and module thermals
02 / SOFTWARE

Runtime & Models

Connect infrastructure behavior to model execution and serving quality.

  • Kernel, operator and graph traces
  • Prefill, decode and KV-cache metrics
  • TTFT, TPOT, throughput and tail latency
  • Batching, scheduling and queue delay
  • Framework, driver and compiler state
03 / SYSTEMS

Clusters & Platforms

Understand coordination across nodes, fabrics, storage and control planes.

  • Node health and topology awareness
  • Collectives and network congestion
  • Storage and checkpoint performance
  • Container, VM and orchestration state
  • Placement, fragmentation and capacity
04 / PHYSICAL

Power, Cooling & Facilities

Relate digital workloads to the physical systems keeping them alive.

  • Rack, row and campus power
  • CDUs, loops, pumps and heat exchangers
  • UPS, switchgear and generator state
  • PUE, WUE and cooling effectiveness
  • Facility alarms and environmental risk
Any silicon, one view

Hardware-agnostic by design.

Drshym reads counters natively across accelerator vendors and correlates them in one schema — no per-vendor dashboards to reconcile.

NVIDIA · CUDAAMD · ROCmIntel · XPUEmerging RISC-V acceleratorsCloud TPU

GPU Utilization Heatmap 32 accelerators

58
98
47
63
79
41
42
90
72
44
61
75
41
96
70
51
40
43
65
64
42
53
43
73
65
41
90
74
45
98
52
78

Tensor Core Precision Mix

FP16 / BF16 — 42%INT8 — 31%FP8 — 19%INT4 — 8%

Node Health 240 nodes

HealthyDegradedDown

Interconnect Bandwidth % of peak

N-A1
N-A2
N-B1
N-B2
N-A1
100
92
61
58
N-A2
92
100
64
55
N-B1
61
64
100
88
N-B2
58
55
88
100

Facility Floor Plan 4 zones

Zone ACompute Hall 1
Zone BCompute Hall 2
Zone CCooling Plant
Zone DPower & Substation

Illustrative data across all panels — sample values, not a live customer environment.

Signal coverage

Every layer. Every timescale. One correlated view.

Drshym captures high-frequency counters, distributed traces, events, topology, environmental state, and historical context without losing their relationships.

Layer
Microseconds
Milliseconds
Seconds
Minutes
Hours
Months
Silicon
Pipeline stalls
Kernel phases
Thermal response
DVFS drift
Fault trends
Aging
Runtime
Operator timing
Queue delay
Batch behavior
Serving mix
Release impact
Model evolution
System
Packet events
Collectives
Node health
Capacity
Failure patterns
Fleet lifecycle
Physical
Sensor reads
Control loops
Rack thermals
Plant efficiency
Energy profile
Facility planning
Digital twin

Understand how every subsystem affects the others.

Drshym maps workload demand through compute, network, electrical, thermal, and facility domains to explain root causes instead of isolated symptoms.

Workload to Silicon

Trace a request through model execution, scheduling, kernels, memory, fabric, and accelerator behavior.

RequestModelSchedulerKernelGPUHBM
APIModelRuntimeGPUHBM

Rack to Utility

Connect GPU demand to power delivery, cooling response, electrical constraints, and campus efficiency.

GPU BoardRackPDUUPSCoolingGrid
RackPDUUPSPlantGrid
Operational loop

Observe, explain, predict, and act.

Drshym supports the full operational cycle, from high-resolution telemetry collection to anomaly explanation, capacity forecasting, and policy-aware recommendations.

01

Observe

Collect counters, traces, logs, events, topology and facility telemetry.

02

Correlate

Join model, software, hardware and physical signals in time and space.

03

Explain

Identify causal chains behind latency, failures, thermal drift and wasted capacity.

04

Predict

Forecast SLA risk, component degradation, power excursions and cooling constraints.

05

Optimize

Recommend changes to placement, power, clocks, batching, cooling and capacity plans.

Incident intelligence

Move from alert storms to root cause.

Drshym converts thousands of low-level signals into explainable incidents with blast radius, evidence, timeline, and recommended response.

CriticalTraining job throughput collapsed 31%Cluster A1712 sec ago
WarningHBM queue depth rising across GPU group 42Rack R1738 sec ago
WarningCDU delta pressure outside expected bandCooling Zone 41 min ago
InfoPower cap policy activated during utility eventCampus West4 min ago
WarningNVLink retry rate exceeds learned baselineNode 42-067 min ago

Explained Incident

Utility curtailment reduced campus power envelope by 6%.
Rack-level policy lowered GPU clocks in three rows.
Collective synchronization amplified straggler delay.
Batch completion time increased and throughput fell 31%.
Recommended action: rebalance placement and isolate throttled racks.

Sample alert feed and incident walkthrough — illustrative, not a live customer environment.

Drshym by Ureka AI

See the whole system.

From the smallest signal to the largest infrastructure decision.