Running the model in the institution
Planned
What it is
The model runs on the institution's own machines. Questions are not sent to a model outside the institution. The inference server has no route out and does not ask for an API key.
The inference server has no route out, so no question reaches a model outside the institution.
Deployment profiles
The installation offers two profiles.
| Profile | Memory reserved | Processors and cards reserved |
|---|---|---|
| Graphics-card profile | 32 GiB | One NVIDIA graphics card |
| Processor profile | 16 GiB | 4 processors |
Note
Those figures are what the installation sets aside. They are not a statement of how many people the server can serve.
Plan your installation
Estimate your load
Read the load estimate at the start of How many people. It is an estimate from model sizes and token rates, not a measured benchmark.
Pick a size
Choose the row in How many people that matches your headcount: a pilot of up to about 300 people, 1,000 people or 5,000 people.
Size the whole installation
Use The whole installation to size the web application, background jobs, database, backups, front door, monitoring and the model machine.
Check model memory and choose a card
Use Memory each model needs and Which graphics card to pick hardware for the models you plan to run.
Keep the work small
Apply the measures in Keeping the work small, so fewer questions reach the large model.
How many people
The tables below are estimates from model sizes and token rates. They are not a measured benchmark. With the same usage, 1,000 people produce about 560 million input tokens and 50 million output tokens a month on the large model. Spread over working hours, that is roughly 1,000 input tokens and 80 output tokens per second on average, and about five times that at peak. A mixture-of-experts model with about 4 billion active parameters handles this on one data-centre graphics card.
| Size | Small models (router, guard, embeddings) | Large model | Server |
|---|---|---|---|
| Pilot, up to about 300 people | Processor: 16 cores, 64 GB of memory (Jan-v1-4B at 4-bit about 3 GB, Gemma 4 E4B about 4 GB) | Gemma 4 26B-A4B on one 24–48 GB graphics card, 4-bit or 8-bit | One graphics-card machine (for example L40S or RTX 6000 class) plus the application servers |
| 1,000 people | The same graphics card as the large model (the small models add about 8 GB) | Gemma 4 26B-A4B or Qwen3.5-35B-A3B, 48 GB graphics card | Two graphics-card machines so one can take over (one is enough for the load) |
| 5,000 people | A separate small graphics card, or two processor machines | Gemma 4 31B, or a 120-billion-parameter mixture-of-experts model, on 80 GB graphics cards | Two to four 80 GB graphics cards (H100 or H200 class), with batching |
The whole installation
Processor counts are physical cores or virtual processors. Disks are fast solid-state drives. The model is optional: without it, the other rows still run on ordinary machines and no graphics card is required.
| Part | Pilot, up to 300 people | 1,000 people | 5,000 people |
|---|---|---|---|
| Web application | 1 × 4 processors, 8 GB | 2 × 8 processors, 16 GB, behind a load balancer | 4 × 8 processors, 16 GB |
| Background jobs | Shares the web machine | 1 × 4 processors, 8 GB | 2 × 8 processors, 16 GB |
| Database | 4 processors, 16 GB, 200 GB | 8 processors, 32 GB, 500 GB, plus a copy that follows it | 16 processors, 64 GB, 1–2 TB, plus a copy |
| Backups and files | 500 GB | 2 TB | 5–10 TB |
| The secure front door | On the web machine | Two small machines (2 processors, 4 GB) | Two small machines |
| Monitoring and logs | 2 processors, 4 GB | 4 processors, 8 GB, 200 GB | 8 processors, 16 GB, 500 GB |
| Model machine | One machine: one 24–48 GB graphics card, 8–16 cores, 64 GB of memory, 1 TB disk | Two machines: one 48 GB graphics card each, 16 cores, 128 GB of memory, 2 TB disk | Two to four machines: 80 GB graphics cards, 32 cores, 256 GB of memory |
Tip
If only the small models run, and only administrators use them, no graphics card is required. One 16-core machine with 64 GB of memory runs Jan-v1-4B, Gemma 4 E4B, EmbeddingGemma and Prompt Guard for up to about 300 people who use the model.
Memory each model needs
These figures are the weights only. Add 30–50% so several people can be served at once. A rough rule is parameters × 0.55 GB at 4-bit and × 1.05 GB at 8-bit. The machine's own memory should be at least twice the graphics-card memory.
| Model | Job | 4-bit | 8-bit | Processor only? |
|---|---|---|---|---|
| Prompt Guard 2 (86 million) | Injection check | under 1 GB | under 1 GB | Yes, and it is fast |
| EmbeddingGemma (308 million) | Search | under 1 GB | about 1 GB | Yes |
| Jan-v1-4B | Choosing a tool | about 3 GB | about 5 GB | Yes (about 10–20 tokens a second for one request) |
| Gemma 4 E4B | Classify and check | about 4–5 GB | about 8 GB | Yes |
| Gemma 4 26B-A4B | Standard large model | about 16 GB | about 28 GB | No |
| Qwen3.5-35B-A3B | Backup large model | about 21 GB | about 37 GB | No |
| Gemma 4 31B | Larger on-prem planner | about 19 GB | about 33 GB | No |
| Jan-v1-4B | 5 GB |
|---|---|
| Gemma 4 E4B | 8 GB |
| Gemma 4 26B-A4B | 28 GB |
| Gemma 4 31B | 33 GB |
| Qwen3.5-35B-A3B | 37 GB |
Which graphics card
| Card | Memory | Power | Fits | Suited to | Note |
|---|---|---|---|---|---|
| NVIDIA L4 | 24 GB | 72 W | The small models and Gemma 4 26B-A4B at 4-bit | A pilot, low power, any 1U server | Slower when many people ask at once |
| RTX 5090 / 4090 | 32 GB / 24 GB | 450–575 W | The same as the L4, and faster | A developer machine or an internal demonstration | The GeForce driver licence does not allow data-centre use, so not for the institution's production servers |
| L40S | 48 GB | 350 W | 26B-A4B at 8-bit, the small models, and room for several people at once | 1,000 people | The usual choice |
| RTX PRO 6000 (Blackwell) | 96 GB | 300–600 W | The 31B model at 8-bit with room, or two models | One card for the large model and the larger planner | Workstation and server editions |
| H100 / H200 | 80 GB / 141 GB | 700 W | The models above, with many people at once | 5,000 or more people | Only when the load needs it |
Warning
The GeForce driver licence does not allow data-centre use. Do not plan RTX 5090 or 4090 cards for the institution's production servers.
Keeping the work small
- Answer with the records first. "My hours this week" is a lookup, not a model call. The aim is that under a quarter of questions reach the large model.
- Offer only the tools that fit the question. The router picks 5–8 tools instead of sending all 40 definitions, which cuts the input by about 60%.
- Keep the start of the prompt stable. The instructions and the tools come first and do not change during the day, so a repeated start can be reused.
- Summarise a long conversation instead of sending the whole history again.
- Prepare summaries overnight, so they are ready at 8:00.
- Each organisation and each person has a limit. When the limit is reached, the assistant switches to the standard model and says so. It does not stop without saying why.
- The larger model is used for project plans and for explanations of who should do the work.
- On the institution's own machines: mixture-of-experts models, 4-bit or 8-bit weights, and one shared graphics card for all of the small models.
FAQ
No. The model runs on the institution's own machines. The inference server has no route out and does not ask for an API key.
Not always. The model is optional, and the rest of the installation runs on ordinary machines. If only the small models run and only administrators use them, one 16-core machine with 64 GB of memory is enough for up to about 300 people who use the model.
No. They are estimates from model sizes and token rates.
No. The GeForce driver licence does not allow data-centre use. These cards suit a developer machine or an internal demonstration.
The assistant switches to the standard model and says so. It does not stop without saying why.
Related pages
Models that stay in the institution
An administrator can run the small assistant models on a machine the institution controls. People see suggestions, never automatic changes.
What assistance can see
A model call receives only the fields a question needs, and no one sees records through assistance that they could not open themselves.