# Running the model in the institution (https://docs.akollo.com/en/ai/on-prem)



<Callout title="Planned">
  This area is planned. It is described here so you can prepare; it is not yet part of the current release.
</Callout>

## What it is [#what-it-is]

The model runs on the institution's own machines. Questions are not sent to a model outside the institution. The
inference server has no route out and does not ask for an API key.

<Mermaid
  title="Data path inside the institution"
  chart="`flowchart LR
A[Employee asks a question] --> B[Akollo application]
B --> C[Inference server]
C --> B
C --x|No route out| D[Model outside the institution]`"
/>

The inference server has no route out, so no question reaches a model outside the institution.

## Deployment profiles [#deployment-profiles]

The installation offers two profiles.

| Profile               | Memory reserved | Processors and cards reserved |
| --------------------- | --------------- | ----------------------------- |
| Graphics-card profile | 32 GiB          | One NVIDIA graphics card      |
| Processor profile     | 16 GiB          | 4 processors                  |

<Callout type="info" title="Note">
  Those figures are what the installation sets aside. They are not a statement of how many people the server can serve.
</Callout>

## Plan your installation [#plan-your-installation]

<Steps>
  <Step>
    ### Estimate your load [#estimate-your-load]

    Read the load estimate at the start of **How many people**. It is an estimate from model sizes and token rates,
    not a measured benchmark.
  </Step>

  <Step>
    ### Pick a size [#pick-a-size]

    Choose the row in **How many people** that matches your headcount: a pilot of up to about 300 people, 1,000 people
    or 5,000 people.
  </Step>

  <Step>
    ### Size the whole installation [#size-the-whole-installation]

    Use **The whole installation** to size the web application, background jobs, database, backups, front door,
    monitoring and the model machine.
  </Step>

  <Step>
    ### Check model memory and choose a card [#check-model-memory-and-choose-a-card]

    Use **Memory each model needs** and **Which graphics card** to pick hardware for the models you plan to run.
  </Step>

  <Step>
    ### Keep the work small [#keep-the-work-small]

    Apply the measures in **Keeping the work small**, so fewer questions reach the large model.
  </Step>
</Steps>

## How many people [#how-many-people]

The tables below are estimates from model sizes and token rates. They are not a measured benchmark. With the same
usage, 1,000 people produce about 560 million input tokens and 50 million output tokens a month on the large model.
Spread over working hours, that is roughly 1,000 input tokens and 80 output tokens per second on average, and about
five times that at peak. A mixture-of-experts model with about 4 billion active parameters handles this on one
data-centre graphics card.

| Size                          | Small models (router, guard, embeddings)                                                     | Large model                                                                               | Server                                                                                      |
| ----------------------------- | -------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| Pilot, up to about 300 people | Processor: 16 cores, 64 GB of memory (Jan-v1-4B at 4-bit about 3 GB, Gemma 4 E4B about 4 GB) | Gemma 4 26B-A4B on one 24–48 GB graphics card, 4-bit or 8-bit                             | One graphics-card machine (for example L40S or RTX 6000 class) plus the application servers |
| 1,000 people                  | The same graphics card as the large model (the small models add about 8 GB)                  | Gemma 4 26B-A4B or Qwen3.5-35B-A3B, 48 GB graphics card                                   | Two graphics-card machines so one can take over (one is enough for the load)                |
| 5,000 people                  | A separate small graphics card, or two processor machines                                    | Gemma 4 31B, or a 120-billion-parameter mixture-of-experts model, on 80 GB graphics cards | Two to four 80 GB graphics cards (H100 or H200 class), with batching                        |

## The whole installation [#the-whole-installation]

Processor counts are physical cores or virtual processors. Disks are fast solid-state drives. The model is optional:
without it, the other rows still run on ordinary machines and no graphics card is required.

| Part                  | Pilot, up to 300 people                                                         | 1,000 people                                                                      | 5,000 people                                                           |
| --------------------- | ------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| Web application       | 1 × 4 processors, 8 GB                                                          | 2 × 8 processors, 16 GB, behind a load balancer                                   | 4 × 8 processors, 16 GB                                                |
| Background jobs       | Shares the web machine                                                          | 1 × 4 processors, 8 GB                                                            | 2 × 8 processors, 16 GB                                                |
| Database              | 4 processors, 16 GB, 200 GB                                                     | 8 processors, 32 GB, 500 GB, plus a copy that follows it                          | 16 processors, 64 GB, 1–2 TB, plus a copy                              |
| Backups and files     | 500 GB                                                                          | 2 TB                                                                              | 5–10 TB                                                                |
| The secure front door | On the web machine                                                              | Two small machines (2 processors, 4 GB)                                           | Two small machines                                                     |
| Monitoring and logs   | 2 processors, 4 GB                                                              | 4 processors, 8 GB, 200 GB                                                        | 8 processors, 16 GB, 500 GB                                            |
| Model machine         | One machine: one 24–48 GB graphics card, 8–16 cores, 64 GB of memory, 1 TB disk | Two machines: one 48 GB graphics card each, 16 cores, 128 GB of memory, 2 TB disk | Two to four machines: 80 GB graphics cards, 32 cores, 256 GB of memory |

<Callout type="idea" title="Tip">
  If only the small models run, and only administrators use them, no graphics card is required. One 16-core machine
  with 64 GB of memory runs Jan-v1-4B, Gemma 4 E4B, EmbeddingGemma and Prompt Guard for up to about 300 people who use
  the model.
</Callout>

## Memory each model needs [#memory-each-model-needs]

These figures are the weights only. Add 30–50% so several people can be served at once. A rough rule is parameters ×
0.55 GB at 4-bit and × 1.05 GB at 8-bit. The machine's own memory should be at least twice the graphics-card memory.

| Model                        | Job                    | 4-bit        | 8-bit       | Processor only?                                   |
| ---------------------------- | ---------------------- | ------------ | ----------- | ------------------------------------------------- |
| Prompt Guard 2 (86 million)  | Injection check        | under 1 GB   | under 1 GB  | Yes, and it is fast                               |
| EmbeddingGemma (308 million) | Search                 | under 1 GB   | about 1 GB  | Yes                                               |
| Jan-v1-4B                    | Choosing a tool        | about 3 GB   | about 5 GB  | Yes (about 10–20 tokens a second for one request) |
| Gemma 4 E4B                  | Classify and check     | about 4–5 GB | about 8 GB  | Yes                                               |
| Gemma 4 26B-A4B              | Standard large model   | about 16 GB  | about 28 GB | No                                                |
| Qwen3.5-35B-A3B              | Backup large model     | about 21 GB  | about 37 GB | No                                                |
| Gemma 4 31B                  | Larger on-prem planner | about 19 GB  | about 33 GB | No                                                |

<BarChart title="Approximate weights at 8-bit" caption="Weights only, from the table above. Add 30–50% to serve several people at once." unit=" GB" data="[{ label: 'Jan-v1-4B', value: 5 }, { label: 'Gemma 4 E4B', value: 8 }, { label: 'Gemma 4 26B-A4B', value: 28 }, { label: 'Gemma 4 31B', value: 33 }, { label: 'Qwen3.5-35B-A3B', value: 37 }]" />

## Which graphics card [#which-graphics-card]

| Card                     | Memory         | Power     | Fits                                                                    | Suited to                                           | Note                                                                                                       |
| ------------------------ | -------------- | --------- | ----------------------------------------------------------------------- | --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| NVIDIA L4                | 24 GB          | 72 W      | The small models and Gemma 4 26B-A4B at 4-bit                           | A pilot, low power, any 1U server                   | Slower when many people ask at once                                                                        |
| RTX 5090 / 4090          | 32 GB / 24 GB  | 450–575 W | The same as the L4, and faster                                          | A developer machine or an internal demonstration    | The GeForce driver licence does not allow data-centre use, so not for the institution's production servers |
| L40S                     | 48 GB          | 350 W     | 26B-A4B at 8-bit, the small models, and room for several people at once | 1,000 people                                        | The usual choice                                                                                           |
| RTX PRO 6000 (Blackwell) | 96 GB          | 300–600 W | The 31B model at 8-bit with room, or two models                         | One card for the large model and the larger planner | Workstation and server editions                                                                            |
| H100 / H200              | 80 GB / 141 GB | 700 W     | The models above, with many people at once                              | 5,000 or more people                                | Only when the load needs it                                                                                |

<Callout type="warn" title="Warning">
  The GeForce driver licence does not allow data-centre use. Do not plan RTX 5090 or 4090 cards for the institution's
  production servers.
</Callout>

## Keeping the work small [#keeping-the-work-small]

<Mermaid
  title="How a question is handled"
  chart="`flowchart TD
A[Question] --> B{Can the records answer it?}
B -->|Yes| C[Lookup, no model call]
B -->|No| D[Router picks 5–8 tools]
D --> E{Limit reached?}
E -->|No| F[Model chosen for the job]
E -->|Yes| G[&#x22;Standard model, and the assistant says so&#x22;]`"
/>

* Answer with the records first. "My hours this week" is a lookup, not a model call. The aim is that under a quarter
  of questions reach the large model.
* Offer only the tools that fit the question. The router picks 5–8 tools instead of sending all 40 definitions, which
  cuts the input by about 60%.
* Keep the start of the prompt stable. The instructions and the tools come first and do not change during the day, so
  a repeated start can be reused.
* Summarise a long conversation instead of sending the whole history again.
* Prepare summaries overnight, so they are ready at 8:00.
* Each organisation and each person has a limit. When the limit is reached, the assistant switches to the standard
  model and says so. It does not stop without saying why.
* The larger model is used for project plans and for explanations of who should do the work.
* On the institution's own machines: mixture-of-experts models, 4-bit or 8-bit weights, and one shared graphics card
  for all of the small models.

## FAQ [#faq]

<Accordions type="single">
  <Accordion title="Does any question leave the institution?">
    No. The model runs on the institution's own machines. The inference server has no route out and does not ask for an
    API key.
  </Accordion>

  <Accordion title="Do we need a graphics card?">
    Not always. The model is optional, and the rest of the installation runs on ordinary machines. If only the small
    models run and only administrators use them, one 16-core machine with 64 GB of memory is enough for up to about 300
    people who use the model.
  </Accordion>

  <Accordion title="Are the sizing figures measured benchmarks?">
    No. They are estimates from model sizes and token rates.
  </Accordion>

  <Accordion title="Can we use RTX 5090 or 4090 cards in production?">
    No. The GeForce driver licence does not allow data-centre use. These cards suit a developer machine or an internal
    demonstration.
  </Accordion>

  <Accordion title="What happens when a person reaches their limit?">
    The assistant switches to the standard model and says so. It does not stop without saying why.
  </Accordion>
</Accordions>

## Related pages [#related-pages]

* [The organisation's own connection](/en/ai/connections)
* [Models that stay in the institution](/en/ai/local-models)
* [What assistance can see](/en/ai/privacy)
