NetSphere

NetSphere LLC · Texas · est. 2026

An AI systems lab that measures first.

We build NetSphere, a sovereign AI platform that runs open-weight models and agents on hardware you own. The lab behind it measures every number it publishes and releases the kernels, drafters and recipes that come out of the work. Every performance figure here was measured on a named machine, and the receipts exist. Spec and price plates say where their numbers came from.

4× RTX PRO 6000 Blackwell 1× DGX Spark 2 servers on custom water Pascal to Blackwell, sm_61 → sm_121

Platform

NetSphere: sovereign AI on hardware you own.

A self-hosted stack for open-weight models and the agents built on them, with no cloud model API in the serving path. Every action that changes a system waits for a person, every GPU is metered, and models are qualified on the machine they will run on. We build it, and the lab runs on it.

Serving

Open models on your GPUs

Models served with vLLM, SGLang or llama.cpp, on anything from Blackwell workstation cards down to Pascal, with launch recipes measured on the machine that runs them.

Agents

An agent whose memory stays home

A tool-using agent for chat, voice and background tasks. Conversations, memory and the knowledge vault live in your own Postgres.

Execution rail

Nothing changes without a person

Every action that changes a system waits for approval. Commands on your hosts run only on a single-use grant that the host itself verifies, never the agent.

Sandboxes

Sealed by default

Code and web browsing run in sealed containers that hold no credentials. Reaching a real system goes through the rail, never around it.

Qualification

From download to verdict, one pipeline

A model goes from download to a verdict on your hardware in one pipeline: fit, serving recipe, speed ladders, concurrency knee and quality gates, with the receipts.

Observability

Every GPU on the record

Power, clocks, temperature and memory for every GPU, plus container logs, kept in Prometheus, Loki and Grafana and shown live on an operations console.

Deployments

Bring NetSphere to your building.

We deploy the platform on your hardware and qualify the models you want to run on it before handover. Nothing leaves the premises, and your operators keep the keys.

Work

Released, with numbers attached.

Everything below runs on the machines in the next section. Each release ships with the run that produced its figures.

Kernels · Pascal

Ternary Bonsai 2 on a GTX 1080 Ti

42.6 → 98.4tok/s, 27B, one 2017 card

Custom llama.cpp kernels for the ternary weight format on Pascal, plus a native multi-token-prediction drafter we trained for it. Chat decode goes from 42.6 tok/s plain to 64 to 98 tok/s depending on the workload, at 32K context on 11 GB of VRAM. MIT.

Measurement · Serving

Serving recipes for open models

N ≥ 3plus a warmup, by rule

Decode, prefill and concurrency ladders for the models people actually run, on Blackwell workstation cards, the DGX Spark and consumer GPUs. Thinking-mode effort, speculative decoding, KV precision and PCIe generation are each measured one variable at a time, and the losing arm is published next to the winner.

Platform · Sovereign

Built on NetSphere

0cloud model APIs in the serving path

The releases above came out of a lab that runs on NetSphere, the platform it develops. It serves the models, runs the agent, gates every change and meters every GPU.

Field plates

One finding per sheet, built from the raw rows.

Field plate 033: GPUs ranked by memory bandwidth per dollar
Plate 033Memory bandwidth per dollar across 48 GPUs, NVIDIA and AMD, priced new, used and by allocation.
Field plate 031: NVIDIA GPUs ranked by memory bandwidth
Plate 031The NVIDIA bandwidth ladder. Decode speed is a bandwidth number before it is anything else.
Field plate 036: Bonsai 2 decode speed as context fills on a GTX 1080 Ti
Plate 036Bonsai 2 on the 1080 Ti as the context fills, with and without the drafter.
Field plate 038: stock versus overclocked RTX PRO 6000 under water
Plate 038Stock against a permanent memory overclock on water: 16 to 19% decode, ECC on, hottest card 51 °C.

The lab

Four machines, nine GPUs, one variable at a time.

Built and maintained in-house, including the water loops. Every host is instrumented down to per-GPU power and clocks, so a number on a plate can be traced to a timestamp on a card.

Primary host4× NVIDIA RTX PRO 6000 Blackwell, 96 GB each · Threadripper PRO 9985WX · 768 GB DDR5Custom water loop with an external pump and radiator tower. Permanent memory and core offset, ECC on, verified by a stock / overclock / stock A/B.
In progressEight-GPU PCIe Gen5 fabric on three Microchip PM50100 switchesA root switch split in two partitions over two leaves, so GPU-to-GPU traffic never crosses the CPU. Bare metal, measured before and after.
Batch and evalNVIDIA DGX Spark, GB10, 128 GB unified memoryStanding resident for evaluation runs and long-context prefill studies.
Second hostRTX 5090 · RTX 3090 · RTX PRO 4500 BlackwellMixed-generation host for the agent platform, voice, and single-card comparisons. Rebuilt on water in 2026.
Legacy laneGeForce GTX 1080 Ti, 11 GB, PascalWhere the Bonsai 2 kernels were written. Kept because the 8 GB crowd is most of the audience.

Method

Rules we do not bend.

Measured, not modelled

A number is published only after it was produced on a named machine. Projections are labelled as projections, on the plate, next to the figure.

One variable at a time

Arms differ in exactly one setting. Thinking mode, speculative decoding and KV precision are held fixed across every body in a comparison.

N of three, plus a warmup

A single run draws noise as a mountain. Every reading is the median of at least three, after a discarded warmup request.

The losing arm is published

Negative results, truncated runs and failed fits go on the plate or in the report. A clean story that hides its failures is not a measurement.

Receipts travel with the claim

Every plate and report names the run directory that produced it. The raw rows are kept and can be requested.

Corrections are public

When a number was wrong, the correction is posted in the same place, with the reason. Several of our best findings started as corrections.

Contact

Talk to the lab.

To bring NetSphere to your hardware, to collaborate on research, or to send a model you want measured.

Email

hello@netsphere.com.ai

X

@net_termina

Where the field plates post first.

Code and weights

GitHub · Hugging Face

Kernels, patches, drafters and model cards.