AI

Running the model in the institution

Ask AI about this page

Planned

This area is planned. It is described here so you can prepare; it is not yet part of the current release.

What it is

The model runs on the institution's own machines. Questions are not sent to a model outside the institution. The inference server has no route out and does not ask for an API key.

Data path inside the institution

The inference server has no route out, so no question reaches a model outside the institution.

Deployment profiles

The installation offers two profiles.

ProfileMemory reservedProcessors and cards reserved
Graphics-card profile32 GiBOne NVIDIA graphics card
Processor profile16 GiB4 processors

Note

Those figures are what the installation sets aside. They are not a statement of how many people the server can serve.

Plan your installation

Estimate your load

Read the load estimate at the start of How many people. It is an estimate from model sizes and token rates, not a measured benchmark.

Pick a size

Choose the row in How many people that matches your headcount: a pilot of up to about 300 people, 1,000 people or 5,000 people.

Size the whole installation

Use The whole installation to size the web application, background jobs, database, backups, front door, monitoring and the model machine.

Check model memory and choose a card

Use Memory each model needs and Which graphics card to pick hardware for the models you plan to run.

Keep the work small

Apply the measures in Keeping the work small, so fewer questions reach the large model.

How many people

The tables below are estimates from model sizes and token rates. They are not a measured benchmark. With the same usage, 1,000 people produce about 560 million input tokens and 50 million output tokens a month on the large model. Spread over working hours, that is roughly 1,000 input tokens and 80 output tokens per second on average, and about five times that at peak. A mixture-of-experts model with about 4 billion active parameters handles this on one data-centre graphics card.

SizeSmall models (router, guard, embeddings)Large modelServer
Pilot, up to about 300 peopleProcessor: 16 cores, 64 GB of memory (Jan-v1-4B at 4-bit about 3 GB, Gemma 4 E4B about 4 GB)Gemma 4 26B-A4B on one 24–48 GB graphics card, 4-bit or 8-bitOne graphics-card machine (for example L40S or RTX 6000 class) plus the application servers
1,000 peopleThe same graphics card as the large model (the small models add about 8 GB)Gemma 4 26B-A4B or Qwen3.5-35B-A3B, 48 GB graphics cardTwo graphics-card machines so one can take over (one is enough for the load)
5,000 peopleA separate small graphics card, or two processor machinesGemma 4 31B, or a 120-billion-parameter mixture-of-experts model, on 80 GB graphics cardsTwo to four 80 GB graphics cards (H100 or H200 class), with batching

The whole installation

Processor counts are physical cores or virtual processors. Disks are fast solid-state drives. The model is optional: without it, the other rows still run on ordinary machines and no graphics card is required.

PartPilot, up to 300 people1,000 people5,000 people
Web application1 × 4 processors, 8 GB2 × 8 processors, 16 GB, behind a load balancer4 × 8 processors, 16 GB
Background jobsShares the web machine1 × 4 processors, 8 GB2 × 8 processors, 16 GB
Database4 processors, 16 GB, 200 GB8 processors, 32 GB, 500 GB, plus a copy that follows it16 processors, 64 GB, 1–2 TB, plus a copy
Backups and files500 GB2 TB5–10 TB
The secure front doorOn the web machineTwo small machines (2 processors, 4 GB)Two small machines
Monitoring and logs2 processors, 4 GB4 processors, 8 GB, 200 GB8 processors, 16 GB, 500 GB
Model machineOne machine: one 24–48 GB graphics card, 8–16 cores, 64 GB of memory, 1 TB diskTwo machines: one 48 GB graphics card each, 16 cores, 128 GB of memory, 2 TB diskTwo to four machines: 80 GB graphics cards, 32 cores, 256 GB of memory

Tip

If only the small models run, and only administrators use them, no graphics card is required. One 16-core machine with 64 GB of memory runs Jan-v1-4B, Gemma 4 E4B, EmbeddingGemma and Prompt Guard for up to about 300 people who use the model.

Memory each model needs

These figures are the weights only. Add 30–50% so several people can be served at once. A rough rule is parameters × 0.55 GB at 4-bit and × 1.05 GB at 8-bit. The machine's own memory should be at least twice the graphics-card memory.

ModelJob4-bit8-bitProcessor only?
Prompt Guard 2 (86 million)Injection checkunder 1 GBunder 1 GBYes, and it is fast
EmbeddingGemma (308 million)Searchunder 1 GBabout 1 GBYes
Jan-v1-4BChoosing a toolabout 3 GBabout 5 GBYes (about 10–20 tokens a second for one request)
Gemma 4 E4BClassify and checkabout 4–5 GBabout 8 GBYes
Gemma 4 26B-A4BStandard large modelabout 16 GBabout 28 GBNo
Qwen3.5-35B-A3BBackup large modelabout 21 GBabout 37 GBNo
Gemma 4 31BLarger on-prem plannerabout 19 GBabout 33 GBNo
Approximate weights at 8-bitWeights only, from the table above. Add 30–50% to serve several people at once.
Jan-v1-4B5 GB
Gemma 4 E4B8 GB
Gemma 4 26B-A4B28 GB
Gemma 4 31B33 GB
Qwen3.5-35B-A3B37 GB

Which graphics card

CardMemoryPowerFitsSuited toNote
NVIDIA L424 GB72 WThe small models and Gemma 4 26B-A4B at 4-bitA pilot, low power, any 1U serverSlower when many people ask at once
RTX 5090 / 409032 GB / 24 GB450–575 WThe same as the L4, and fasterA developer machine or an internal demonstrationThe GeForce driver licence does not allow data-centre use, so not for the institution's production servers
L40S48 GB350 W26B-A4B at 8-bit, the small models, and room for several people at once1,000 peopleThe usual choice
RTX PRO 6000 (Blackwell)96 GB300–600 WThe 31B model at 8-bit with room, or two modelsOne card for the large model and the larger plannerWorkstation and server editions
H100 / H20080 GB / 141 GB700 WThe models above, with many people at once5,000 or more peopleOnly when the load needs it

Warning

The GeForce driver licence does not allow data-centre use. Do not plan RTX 5090 or 4090 cards for the institution's production servers.

Keeping the work small

How a question is handled
  • Answer with the records first. "My hours this week" is a lookup, not a model call. The aim is that under a quarter of questions reach the large model.
  • Offer only the tools that fit the question. The router picks 5–8 tools instead of sending all 40 definitions, which cuts the input by about 60%.
  • Keep the start of the prompt stable. The instructions and the tools come first and do not change during the day, so a repeated start can be reused.
  • Summarise a long conversation instead of sending the whole history again.
  • Prepare summaries overnight, so they are ready at 8:00.
  • Each organisation and each person has a limit. When the limit is reached, the assistant switches to the standard model and says so. It does not stop without saying why.
  • The larger model is used for project plans and for explanations of who should do the work.
  • On the institution's own machines: mixture-of-experts models, 4-bit or 8-bit weights, and one shared graphics card for all of the small models.

FAQ

On this page