Coming soon — closed beta

Tensor
Relay

Clustered AI Inference

Join with peers to pool GPU resources for coding, game development, creative projects, research, agentic AI, and much more.

Join the betaFree during beta · Windows & Linux

One cluster from many machines

Tensor Relay splits a model's layers across the machines in your cluster. Each peer contributes GPU memory and compute — together they run what none could alone.

Your hardware, your models

Run open-weight models on hardware you already own. No cloud bill, no API keys, no data leaving your cluster.

Ships on Steam

A desktop app for Windows and Linux, installed and updated through Steam. Join a cluster with friends in a couple of clicks.

Powerful frontier models produce stronger reasoning, richer code assistance, and more precise creative results because they carry more parameters, context, and learned capability. That extra quality usually means larger model weights and large VRAM requirements.

Tensor Relay makes those larger models practical by splitting work across participating machines. Each peer contributes GPU memory and compute, so a cluster can run models that would be difficult or impossible for average users.

Pool hardware with friends or public peers, and unlock access to high quality inference from home.

Tell us what you'd run it on — invites go out in waves.