Open models on your GPUs
Models served with vLLM, SGLang or llama.cpp, on anything from Blackwell workstation cards down to Pascal, with launch recipes measured on the machine that runs them.
NetSphere LLC · Texas · est. 2026
We build NetSphere, a sovereign AI platform that runs open-weight models and agents on hardware you own. The lab behind it measures every number it publishes and releases the kernels, drafters and recipes that come out of the work. Every performance figure here was measured on a named machine, and the receipts exist. Spec and price plates say where their numbers came from.
Platform
A self-hosted stack for open-weight models and the agents built on them, with no cloud model API in the serving path. Every action that changes a system waits for a person, every GPU is metered, and models are qualified on the machine they will run on. We build it, and the lab runs on it.
Models served with vLLM, SGLang or llama.cpp, on anything from Blackwell workstation cards down to Pascal, with launch recipes measured on the machine that runs them.
A tool-using agent for chat, voice and background tasks. Conversations, memory and the knowledge vault live in your own Postgres.
Every action that changes a system waits for approval. Commands on your hosts run only on a single-use grant that the host itself verifies, never the agent.
Code and web browsing run in sealed containers that hold no credentials. Reaching a real system goes through the rail, never around it.
A model goes from download to a verdict on your hardware in one pipeline: fit, serving recipe, speed ladders, concurrency knee and quality gates, with the receipts.
Power, clocks, temperature and memory for every GPU, plus container logs, kept in Prometheus, Loki and Grafana and shown live on an operations console.
We deploy the platform on your hardware and qualify the models you want to run on it before handover. Nothing leaves the premises, and your operators keep the keys.
Work
Everything below runs on the machines in the next section. Each release ships with the run that produced its figures.
42.6 → 98.4tok/s, 27B, one 2017 card
Custom llama.cpp kernels for the ternary weight format on Pascal, plus a native multi-token-prediction drafter we trained for it. Chat decode goes from 42.6 tok/s plain to 64 to 98 tok/s depending on the workload, at 32K context on 11 GB of VRAM. MIT.
N ≥ 3plus a warmup, by rule
Decode, prefill and concurrency ladders for the models people actually run, on Blackwell workstation cards, the DGX Spark and consumer GPUs. Thinking-mode effort, speculative decoding, KV precision and PCIe generation are each measured one variable at a time, and the losing arm is published next to the winner.
0cloud model APIs in the serving path
The releases above came out of a lab that runs on NetSphere, the platform it develops. It serves the models, runs the agent, gates every change and meters every GPU.
Field plates




The lab
Built and maintained in-house, including the water loops. Every host is instrumented down to per-GPU power and clocks, so a number on a plate can be traced to a timestamp on a card.
Method
A number is published only after it was produced on a named machine. Projections are labelled as projections, on the plate, next to the figure.
Arms differ in exactly one setting. Thinking mode, speculative decoding and KV precision are held fixed across every body in a comparison.
A single run draws noise as a mountain. Every reading is the median of at least three, after a discarded warmup request.
Negative results, truncated runs and failed fits go on the plate or in the report. A clean story that hides its failures is not a measurement.
Every plate and report names the run directory that produced it. The raw rows are kept and can be requested.
When a number was wrong, the correction is posted in the same place, with the reason. Several of our best findings started as corrections.
Contact
To bring NetSphere to your hardware, to collaborate on research, or to send a model you want measured.
hello@netsphere.com.ai