We're building the future of compute.
Software and hardware, co-designed.
Software today — any workload on any chip.
Co-located heterogeneous hardware coming next.
For two decades, advanced compute meant two chips. That era is ending — deciding what runs where stops being craft. It becomes infrastructure.
Within a few years a single data center — often a single rack — will hold five, ten, or more classes of silicon: CPU and GPU joined by TPUs, LPUs, wafer-scale engines, quantum and more — each fastest at a different kind of work, each with its own cost, energy and error profile.
Choosing the chip and the algorithm by hand works for two chips. It breaks at three and more. Deciding what runs where stops being a matter of craft and becomes infrastructure. That is the layer ORIQX occupies — it decides how a workload computes, above the schedulers that place jobs and beneath the tools that call them. It pays off today on your CPU and GPU fleet, and it is the only practical way to program the massively heterogeneous racks coming next.
ORIQX is that layer: the infrastructure for the future of compute.
One workload, decomposed across chips.
A single model on one chip does everything adequately — and little of it well. ORIQX splits the workload into compute blocks and maps each to the chip that runs it best.
Without ORIQX
- Manual tuning for a specific hardware target
- Often under-optimized
- Not tailored to your specific infrastructure
With ORIQX
- Optimized — tailored to your workload × infrastructure
- No manual work — no vendor-specific integration
- Always dynamic — decouple your hardware from your software
Your code, any language — every block on its best-fit chip.
Stick to your own code — Julia, C, C++, Python, or any other language. ORIQX understands your computational intent, decomposes it into compute blocks, prices each across the reachable silicon, and runs each where it fits — no vendor-specific code.
Twenty-eight classes of silicon, one program.
ORIQX reasons over twenty-eight classes of processor as a single parameter space — each a vector of throughput, memory, precision, cost and energy, all modeled by ORIQX's HES (Heterogeneous Execution Simulator) from published specifications. No class wins everywhere.
All 28 simulated in HES — tiered by how far each is validated:
Superscalar + SIMD/vector
Wins on
Control flow, sparse work, glue
Loses on
Dense throughput at scale
Massively parallel SIMT
Wins on
Dense linear algebra, batch
Loses on
Branchy/sparse, tiny-batch latency
Qubits — gate & analog
Wins on
Structured exponential subspaces
Loses on
Arithmetic-heavy, unstructured work
Statevector & tensor-network simulation
Wins on
Small-qubit exact validation
Loses on
Qubit counts beyond exact simulation
Systolic tensor arrays
Wins on
Large dense matmul, training
Loses on
Irregular/sparse & control flow
Physical stochastic dynamics
Wins on
Sampling, probabilistic inference
Loses on
Deterministic exact arithmetic
Tiled MAC + scratchpads
Wins on
Edge / client inference
Loses on
Datacenter-scale training
Deterministic dataflow, SRAM-only
Wins on
Low-latency token generation
Loses on
Large working sets (SRAM-bound)
Fine-grained MIMD
Wins on
Fine-grained irregular parallelism
Loses on
Dense regular matmul vs GPUs
Wafer-scale mesh
Wins on
Extreme-bandwidth stencils / ML
Loses on
Small jobs; cost per node
Reconfigurable dataflow
Wins on
Fused pipelines, large models
Loses on
Small or highly-branchy work
Fixed-function tensor
Wins on
Hyperscale training / inference
Loses on
Anything off its fixed function
Digital in-memory MAC
Wins on
Dense low-power inference
Loses on
High-precision work or training
Compute in / near DRAM
Wins on
Memory-bound streaming ops
Loses on
Compute-bound dense math
Spatial reconfigurable logic
Wins on
Streaming, custom precision
Loses on
Peak dense FLOPs; long dev cycles
Interferometric linear optics
Wins on
Ultra-low-energy matmul; interconnect
Loses on
Nonlinear ops; precision (early)
Event-driven spiking cores
Wins on
Sparse temporal, ultra-low power
Loses on
Dense numeric throughput
Quantum annealing / analog Hamiltonian
Wins on
Combinatorial optimization
Loses on
General-purpose compute
Analog in-memory crossbar
Wins on
Ultra-low-energy analog matmul
Loses on
High precision; training
Classical Ising / annealing silicon
Wins on
Combinatorial QUBO optimization
Loses on
General-purpose compute
VLIW fixed-point signal cores
Wins on
Real-time FFT / filters / signal
Loses on
Large dense training
Long-vector HPC engine
Wins on
Bandwidth-bound FP64 HPC
Loses on
Small / irregular work
Networking / storage offload SoC
Wins on
Data movement, security, storage offload
Loses on
Dense compute / FLOPs
In-storage near-NAND compute
Wins on
Near-data scan / filter / compress
Loses on
Compute-bound math
Superconducting single-flux-quantum
Wins on
Ultra-fast, ultra-low-energy logic (cryo)
Loses on
Room-temp deployment; maturity
Rydberg analog Hamiltonian
Wins on
Analog quantum optimization / simulation
Loses on
Arithmetic-heavy classical work
Gaussian boson sampling / MBQC
Wins on
Sampling-class quantum problems
Loses on
General gate computation (early)
Homomorphic-encryption ASIC (NTT)
Wins on
Compute on encrypted data
Loses on
General compute; maturity
Four classes are exercised in production today; eighteen more are modeled and simulated in HES from published specs — four of those also cross-checked against the vendor's own simulator — and six are planned. ORIQX does not claim to run on all twenty-eight; it simulates all twenty-eight, and prices them all.
The verdict follows the computation.
Every workload is a sequence of computational blocks — matmuls, decodes, solves, samples. ORIQX uses novel methods to score each block against every chip — weighing each chip's performance benefit against the cost of moving the data to it — and runs it where it fits best; the same block flips chips as size, precision and data location change. Below: the per-block verdicts, then how they compose into whole workloads.
compute-bound dense tensor math
re-reads weights per token — SRAM-resident wins
near-memory, bandwidth-bound
irregular, SRAM-resident
on-wafer residency beats HBM round-trips
memory-bandwidth-bound
physical sampling is the primitive
compute-bound at full precision — an honest 1×
the interconnect, not the chip, decides
adding a GPU loses to the transfer cost
branchy, sparse, sequential
Compose a workload — then simulate its placement
Drag blocks in (or tap to add), set each block's size, precision and where its data lives, then Simulate — ORIQX decomposes the workload and places each block on its best-fit chip. Change the objective and watch the placement flip.
Your workload — a graph; draw edges left → right
Drag a block's right dot to another box to connect; drag nodes to arrange. Each node shows a sign — hover for the method. Press Simulate to place each block on its best-fit chip; bright edges then show data crossing chips. Try flipping a block's job size S → L: small work stays on the host (the offload doesn't pay), large work moves to an accelerator.
ORIQX hasn't placed this graph yet.
The chip each block runs on is the result of the placement — press Simulate to solve it and compare against a CPU+GPU node.
Always on the Pareto frontier.
ORIQX predicts time, cost, accuracy and energy across the reachable model-hardware space and places you on the Pareto frontier — and keeps you there as models and hardware drift.
What comes next.
One program, written once, placed by ORIQX across multiple co-located racks: every class of silicon on a priced fabric, joined by specialized interconnects. From here: what HES says the rack yields against a CPU+GPU node, how we derive it, and the same rack at datacenter scale.
The co-located rack vs the status quo.
HES simulates the same decomposed workloads on the ORIQX rack and on a conventional CPU+GPU node — speed-up, power saving, and the interconnect that makes both possible.
Recognized by
