← Portfolio · Portfolio case · AI infrastructure, edge inference · United States

OpenInfer: data-centre-scale AI on the hardware you already have

Seed investor · openinfer.io · All companies

OpenInfer brings data-centre-scale AI capabilities to low-end compute and edge devices, and lets fleets of GPUs serve far more tokens inside SLA by scheduling requests across pooled models instead of pinning one model per GPU.

Why we invested

As intelligence gets cheaper, the market expands and supplier lock-in shrinks; whoever makes existing silicon do more wins. OpenInfer's thesis, that the binding constraint in agentic inference is scheduling, not silicon, is an expertise moat that sits exactly where the value is moving.

Milestones

  • October 2025. Joined the Intel Partner Alliance and Microsoft for Startups (announcement).
  • May 2026. OpenClaw beta: vertical disaggregation across GPU and CPU silicon, 50% more inference capacity on an AWS g6e.16xlarge with no additional hardware.
  • July 2026. Fleet benchmark: two identical four-GPU fleets, same models, same traffic; pooling delivered 53.9k tokens inside SLA against 21.4k, zero rejected requests against 1,725, peak utilisation up from 21.5% to 43.5%.
  • August 2026. Intel published OpenInfer's Xeon 6 SoC benchmark in its own Data Center Content Library: 3x faster AI inference on "just a CPU".

Milestones are compiled from the company's own announcements and our Dispatch newsletter; figures are the company's.