Your enterprise’s own model cloud — one server, or a group, sized to your fleet of minds.
In a Stera deployment, the computers your minds live on stay ordinary — no mind does model inference on its own machine. All of it goes to one place: the Forge, your organization’s AI server. It does for your Scintillas exactly what the cloud model providers do for the world — inference, and the computing jobs around it — except it stands inside your walls, serves only your fleet, and sends nothing anywhere.
That makes sizing refreshingly simple. Two dials set the machine: how many minds you run, and how large a local model you want them thinking with. And because minds work in sittings — bursts of thought, not constant streams — a Forge serves far more minds than a naive per-seat calculation suggests. With Scintilla V3, the same server also runs the fleet’s compile jobs: each mind’s knowledge compiling into its own weights, typically in the quiet hours.
Classes, not SKUs — every Forge is sized and quoted with you. The cost driver is the model class you choose, never your headcount.
Single-GPU server · small-model class
A workstation-class server with one strong GPU. Runs small, capable local models — the class that handles routine cognition well — serving a first fleet of role-minds, with V3 compiles running off-hours on the same box. The entry rung for a department or a small company.
Multi-GPU server · mid-model class
Several GPUs in one chassis. Larger local models — or several small ones side by side — with headroom for concurrent sittings and compile jobs without contention. Serves a role fleet across many departments; the workhorse rung for a mid-sized organization.
Server group · large-model class
Dedicated inference nodes plus a compile node, as a group. Frontier-scale local models if you want them, hundreds of minds, redundancy and failover, and an air-gapped variant that runs the same architecture unchanged. The rung for large enterprises, healthcare networks, and government.
A Forge does not have to be all or nothing. Many organizations will run routine cognition on their own server and still let a mind escalate its hardest reasoning to a rented frontier model — by the mind’s own judgment of what a task needs, with the organization’s policy setting the boundary. Others, in regulated environments, will run everything local from day one. Both are first-class configurations, and moving between them changes no mind’s knowledge — the accumulation lives on your machines either way.
And the ladder is climbable in both directions: start on rented muscle with no Forge at all, add a Forge One when the fleet justifies it, grow to a Group as roles multiply. Every rung is a capital expense replacing rent — and the models a Forge runs keep improving, because open-weight engines keep improving and a Forge is never married to one.
Tell us two things — roughly how many roles you want minds for, and whether your compliance posture allows rented escalation or requires everything inside the walls. We will come back with a concrete Forge configuration, its model class, and what it will cost to run — in electricity, not tokens.