Skip to content

We're building the future of compute.

Software and hardware, co-designed.

Software today — any workload on any chip.

Co-located heterogeneous hardware coming next.

The shift

For two decades, advanced compute meant two chips. That era is ending — deciding what runs where stops being craft. It becomes infrastructure.

Within a few years a single data center — often a single rack — will hold five, ten, or more classes of silicon: CPU and GPU joined by TPUs, LPUs, wafer-scale engines, quantum and more — each fastest at a different kind of work, each with its own cost, energy and error profile.

Choosing the chip and the algorithm by hand works for two chips. It breaks at three and more. Deciding what runs where stops being a matter of craft and becomes infrastructure. That is the layer ORIQX occupies — it decides how a workload computes, above the schedulers that place jobs and beneath the tools that call them. It pays off today on your CPU and GPU fleet, and it is the only practical way to program the massively heterogeneous racks coming next.

ORIQX is that layer: the infrastructure for the future of compute.

The mechanism

One workload, decomposed across chips.

A single model on one chip does everything adequately — and little of it well. ORIQX splits the workload into compute blocks and maps each to the chip that runs it best.

Without ORIQX

WORKLOADCPUGPUTPUNPUQPU
  • Manual tuning for a specific hardware target
  • Often under-optimized
  • Not tailored to your specific infrastructure

With ORIQX

CPUGPUTPULPUWSERDUIPUPIMNPUAI ASICAIPUTSU · PPUQPUIsing machinePhotonicNeuromorphicDSPAIMCWORKLOAD
Optimize your compute today
Prepare for heterogeneous advantage tomorrow
  • Optimized — tailored to your workload × infrastructure
  • No manual work — no vendor-specific integration
  • Always dynamic — decouple your hardware from your software
Bring your own code

Your code, any language — every block on its best-fit chip.

Stick to your own code — Julia, C, C++, Python, or any other language. ORIQX understands your computational intent, decomposes it into compute blocks, prices each across the reachable silicon, and runs each where it fits — no vendor-specific code.

YOUR WORKLOAD IS MANY METHODS → ORIQX DECOMPOSES → MAPS EACH BLOCK TO THE CHIP CLASS THAT WINSDense matmulHOVER A METHOD →WHAT YOUR CODE COMPUTESInferencePrefill GEMMAttention decodeKV-cacheMoE routingSpeculative decodeSimulationDFTHF / SCFCCSD(T)MDFEMCFDLattice QCDOptimizationMonte CarloMCMCSimulated annealingQUBO / IsingLP / MILPConjugate gradientQuantizationINT8 / INT4FP8GPTQAWQSmoothQuantQATTrainingGEMM fwd/bwdAll-reduceAdam stepFlashAttentionTensor parallelSignal & analysisFFTFIR / IIRSVDEigensolvePCALU / QRCryptanalysisAESSHA-256RSA modexpKyber (NTT)Shor / GroverGraph & retrievalBFS / PageRankSpGEMMANN (HNSW)Hash joinGNNTPUGoogleGPUNVIDIA · AMD · IntelLPUGroqWSECerebrasRDUSambaNovaIPUGraphcorePIMSamsung · SK hynix · UPMEMNPUQualcomm · Apple · AMD · IntelHuaweiAI ASICAWS · Google · Intel GaudiMeta MTIA · Microsoft MaiaAIPUTenstorrent · d-Matrix · HailoTSU · PPUExtropic · Normal ComputingQPUIBM · IonQ · QuantinuumRigetti · GoogleIsing machineFujitsu · Toshiba · HitachiD-WavePhotonicLightmatter · LightelligenceCelestial AINeuromorphicIntel Loihi · BrainChipSynSense · IBMDSPTI · Qualcomm · Cadence · CEVAAnalog DevicesAIMCMythic · EnCharge · Axelera
Dense matmul · dense linear algebra → TPUDense eigensolve · high-precision FP64 → GPUAttention decode · SRAM-bound · low latency → LPUAll-reduce · bandwidth-bound collective → WSESparse matmul · sparse · conditional → RDUGraph traversal · irregular control flow → IPUStreaming gather · memory-bound → PIMConvolution · tiled MAC · edge → NPUBatched GEMM · fixed-function tensor → AI ASICINT8 matmul · int8 · dense → AIPUMCMC sampling · stochastic sampling → TSU · PPUAmplitude estimation · quantum · exponential state → QPUQUBO anneal · combinatorial search → Ising machineOptical GEMM · low-energy linear → PhotonicSpiking inference · event-driven · sparse → NeuromorphicFFT / filters · fixed-point signal → DSPAnalog matmul · in-memory MAC → AIMC
The hardware catalog

Twenty-eight classes of silicon, one program.

ORIQX reasons over twenty-eight classes of processor as a single parameter space — each a vector of throughput, memory, precision, cost and energy, all modeled by ORIQX's HES (Heterogeneous Execution Simulator) from published specifications. No class wins everywhere.

All 28 simulated in HES — tiered by how far each is validated:

Exercised4
Modeled / simulated18
Planned6
CPU

Superscalar + SIMD/vector

Wins on

Control flow, sparse work, glue

Loses on

Dense throughput at scale

GPU

Massively parallel SIMT

Wins on

Dense linear algebra, batch

Loses on

Branchy/sparse, tiny-batch latency

QPU (gate)

Qubits — gate & analog

Wins on

Structured exponential subspaces

Loses on

Arithmetic-heavy, unstructured work

Quantum simulator

Statevector & tensor-network simulation

Wins on

Small-qubit exact validation

Loses on

Qubit counts beyond exact simulation

TPU

Systolic tensor arrays

Wins on

Large dense matmul, training

Loses on

Irregular/sparse & control flow

TSU · PPU

Physical stochastic dynamics

Wins on

Sampling, probabilistic inference

Loses on

Deterministic exact arithmetic

NPU

Tiled MAC + scratchpads

Wins on

Edge / client inference

Loses on

Datacenter-scale training

LPU

Deterministic dataflow, SRAM-only

Wins on

Low-latency token generation

Loses on

Large working sets (SRAM-bound)

IPU

Fine-grained MIMD

Wins on

Fine-grained irregular parallelism

Loses on

Dense regular matmul vs GPUs

WSE

Wafer-scale mesh

Wins on

Extreme-bandwidth stencils / ML

Loses on

Small jobs; cost per node

RDU

Reconfigurable dataflow

Wins on

Fused pipelines, large models

Loses on

Small or highly-branchy work

AI ASIC

Fixed-function tensor

Wins on

Hyperscale training / inference

Loses on

Anything off its fixed function

AIPU

Digital in-memory MAC

Wins on

Dense low-power inference

Loses on

High-precision work or training

PIM

Compute in / near DRAM

Wins on

Memory-bound streaming ops

Loses on

Compute-bound dense math

FPGA

Spatial reconfigurable logic

Wins on

Streaming, custom precision

Loses on

Peak dense FLOPs; long dev cycles

Photonic

Interferometric linear optics

Wins on

Ultra-low-energy matmul; interconnect

Loses on

Nonlinear ops; precision (early)

Neuromorphic

Event-driven spiking cores

Wins on

Sparse temporal, ultra-low power

Loses on

Dense numeric throughput

QPU (analog)

Quantum annealing / analog Hamiltonian

Wins on

Combinatorial optimization

Loses on

General-purpose compute

AIMC

Analog in-memory crossbar

Wins on

Ultra-low-energy analog matmul

Loses on

High precision; training

Ising machine

Classical Ising / annealing silicon

Wins on

Combinatorial QUBO optimization

Loses on

General-purpose compute

DSP

VLIW fixed-point signal cores

Wins on

Real-time FFT / filters / signal

Loses on

Large dense training

Vector

Long-vector HPC engine

Wins on

Bandwidth-bound FP64 HPC

Loses on

Small / irregular work

DPU

Networking / storage offload SoC

Wins on

Data movement, security, storage offload

Loses on

Dense compute / FLOPs

Computational storage

In-storage near-NAND compute

Wins on

Near-data scan / filter / compress

Loses on

Compute-bound math

SFQ

Superconducting single-flux-quantum

Wins on

Ultra-fast, ultra-low-energy logic (cryo)

Loses on

Room-temp deployment; maturity

Neutral-atom QPU

Rydberg analog Hamiltonian

Wins on

Analog quantum optimization / simulation

Loses on

Arithmetic-heavy classical work

Photonic quantum

Gaussian boson sampling / MBQC

Wins on

Sampling-class quantum problems

Loses on

General gate computation (early)

FHE accelerator

Homomorphic-encryption ASIC (NTT)

Wins on

Compute on encrypted data

Loses on

General compute; maturity

Four classes are exercised in production today; eighteen more are modeled and simulated in HES from published specs — four of those also cross-checked against the vendor's own simulator — and six are planned. ORIQX does not claim to run on all twenty-eight; it simulates all twenty-eight, and prices them all.

Block by block

The verdict follows the computation.

Every workload is a sequence of computational blocks — matmuls, decodes, solves, samples. ORIQX uses novel methods to score each block against every chip — weighing each chip's performance benefit against the cost of moving the data to it — and runs it where it fits best; the same block flips chips as size, precision and data location change. Below: the per-block verdicts, then how they compose into whole workloads.

Dense matmul / GEMMGPU · TPU

compute-bound dense tensor math

Attention decodeLPU

re-reads weights per token — SRAM-resident wins

KV-cache & retrievalPIM

near-memory, bandwidth-bound

Sparse SpMVLPU

irregular, SRAM-resident

DFT / SCF working set (~40 GB)WSE

on-wafer residency beats HBM round-trips

Stencil / molecular dynamicsPIM

memory-bandwidth-bound

MCMC / QUBO samplingTSU · annealer

physical sampling is the primitive

FP64 triples / correctionGPU

compute-bound at full precision — an honest 1×

Collective all-reduceOptical fabric

the interconnect, not the chip, decides

Small dense eigensolveCPU

adding a GPU loses to the transfer cost

Control flow & glueCPU

branchy, sparse, sequential

Compose a workload — then simulate its placement

Drag blocks in (or tap to add), set each block's size, precision and where its data lives, then Simulate — ORIQX decomposes the workload and places each block on its best-fit chip. Change the objective and watch the placement flip.

Start from

Blocks — drag onto the canvas

Your workload — a graph; draw edges left → right

100%
click to remove edgeclick to remove edgeclick to remove edgeclick to remove edgeclick to remove edge
Prefill GEMM
Prefill GEMMfixed-function tensor
Attention decode
Attention decodeSRAM-bound · low latency
KV-cache
KV-cachememory-bound
MoE routing
MoE routingsparse · conditional
Speculative decode
Speculative decodeSRAM-bound · low latency
All-reduce
All-reducefixed-function tensor

Drag a block's right dot to another box to connect; drag nodes to arrange. Each node shows a sign — hover for the method. Press Simulate to place each block on its best-fit chip; bright edges then show data crossing chips. Try flipping a block's job size S → L: small work stays on the host (the offload doesn't pay), large work moves to an accelerator.

Tuning: Prefill GEMM

ORIQX hasn't placed this graph yet.

The chip each block runs on is the result of the placement — press Simulate to solve it and compare against a CPU+GPU node.

Continuously optimized

Always on the Pareto frontier.

ORIQX predicts time, cost, accuracy and energy across the reachable model-hardware space and places you on the Pareto frontier — and keeps you there as models and hardware drift.

TimeCostAccuracyEnergyFastestBalancedCheapest

Higher is better on every spoke · hover a pick to read its values

Software + hardware

What comes next.

One program, written once, placed by ORIQX across multiple co-located racks: every class of silicon on a priced fabric, joined by specialized interconnects. From here: what HES says the rack yields against a CPU+GPU node, how we derive it, and the same rack at datacenter scale.

ORIQX SOFTWAREANY WORKLOAD, ANY CHIPORIQX HARDWAREHETEROGENEOUS SILICONCUSTOMIZED RACKSWORKLOAD OPTIMIZEDORIQXSPECIALIZED SILICONCPUGPUTPUNPULPUIPUAI ASICORIQXSPECIALIZED SILICONCPUWSEPIMFPGANeuromorphicAIMCIsing machineSPECIALIZED INTERCONNECTQPUPLACES EACH BLOCK
HES results

The co-located rack vs the status quo.

HES simulates the same decomposed workloads on the ORIQX rack and on a conventional CPU+GPU node — speed-up, power saving, and the interconnect that makes both possible.

vs a CPU+GPU node
Makespan speed-up, by workloadORIQX rack ÷ CPU+GPU node
Agentic AI
2.5×
LLM inference serving
8.1×
Frontier-model training
4.8×
Computational chemistry
2.7×
Quantitative finance
2.2×
ModeledHES-modeled (makespan) — architectural, not measured; each workload is a decomposed pipeline placed across the rack, against the same pipeline on a CPU+GPU node.

Recognized by

Innosuisse
AWS Activate
CSCS — Swiss National Supercomputing Centre