SC26, November 15-20, 2026, Chicago, Illinois, USA.
Ankush Jain*, Sheng Jiang*, Charles D. Cranor*, Qing Zheng† Brian Atkinson†, George Amvrosiadis*, Gary A. Grider†
*Carnegie Mellon University
†Los Alamos National Laboratory
The bulk-synchronous parallel (BSP) paradigm powers critical computational workloads from scientific simulations to large-scale machine learning training. Running on tens of thousands of processors, these systems are uniquely sensitive to performance variations that current observability approaches struggle to diagnose—post-hoc analysis of massive traces introduces latency and rigidity that obstructs timely diagnosis. To address these challenges, we present ORCA, a system for always-on, steerable BSP observability. ORCA enables operators to materialize relevant views from running applications on demand and intervene in response, via a tree-based overlay providing coordinated collection, in-situ SQL analytics, and timestep-consistent control. ORCA accelerates observability workflows across the spectrum. On a real AMR code at 4096 ranks, ORCA collects equivalent traces at ∼2% overhead (vs 56–81% for tracing tools), enables up to 7275× faster queries, and 99.999% in-situ data reduction. In an LLM-driven evaluation of steerability, a performance anomaly that took us months to diagnose manually is localized in 27 minutes.
FULL TR: pdf