Research & white papers.

Two published papers, and the benchmark results behind them. Real sites, real distance between them, no synthetic tests.

Research

The work, in the open.

Our claims are published, cited and available in full. Read the papers before you believe the marketing.

PublishedarXiv:2608.06188 · 2026

Routing LLM Inference to the Cleanest Grid in Real Time

A. Bernhard, A. B. Yardimci — Solyx AI

Where a request runs decides how much carbon it burns, because the grid is far dirtier in some places and at some hours than others. We showed live, on real GPUs across several regions, that inference can be steered toward the cleaner grid without a single failed request. Replaying a full year of grid data, that steering cuts the emissions of the same workload by roughly half.
Read on arXiv →
PublishedarXiv:2606.15050 · 2026

Solyx AI Grid: Hardware-Telemetry-Aware Routing Across Geographically Distributed GPU Clusters

A. Bernhard, N. Katla — Solyx AI

Ten live signals — from the GPUs, the software and the network between sites — decide where each request runs. Measured against the standard approach on real hardware in three locations: about 1.6 to 1.75 times the useful work from the same fleet, a quarter off the slowest responses, and recovery from a failed site in about a second instead of four.
Read on arXiv →
1.56–1.75×
More usable capacity from the same GPU fleet
At tier-2 SLO, across all eight workload classes tested
99.57%
Long-prompt success rate vs 67.89% for round-robin
0.43% traffic reached misconfigured endpoint · round-robin sends 32.11%
3.2×
Faster failover than round-robin
1,247ms vs 4,226ms P99 reroute · broken site isolated automatically
0.2ms
Routing overhead per request
Minimal overhead. Maximum intelligence.
▼ RESEARCH v2 — sandbox. Head-to-head table, methodology and limitations.

Head to head

Signal-aware routing vs round-robin, on the same hardware.

Every figure below is measured, not modeled. Identical fleet, identical workload, identical SLO target — the only variable is the placement decision.

Metric
Round-robin
Solyx AI Grid
Long-prompt success rate
67.89%
99.57%
Traffic reaching a misconfigured endpoint
32.11%
0.43%
P99 reroute after endpoint failure
4,226 ms
1,247 ms
Usable capacity at the same SLO
1.00×
1.56–1.75×
Control-plane routing overhead
0.2 ms

How it was measured

Two independent benchmark campaigns against a 216-cell SLO matrix, spanning the major inference workload classes. The control plane reads ten routing signals — four application-layer, four hardware, two network — and sits above unmodified inference engines, emitting standard Envoy xDS endpoint weights. The comparison baseline is a real production pressure-based router, not uniform placement.

Want to see the full methodology?

Let's Talk →