# Run your own AI model (https://docs.akollo.com/en/ai/local-setup)



Akollo's AI features run on the models Akollo provides until an admin connects something else. This guide is for
the case where you want the model on hardware you control: a laptop for a trial, a workstation, or a server in your
own network. You install a model runtime, start it as an OpenAI-compatible server, and add it in Akollo as an
**Own model server** connection, directly or through the **AI Connector**.

## Why run your own model [#why-run-your-own-model]

* **Data stays with you.** Questions and answers go to a machine your organisation runs, not to an outside AI
  provider.
* **You choose the model.** You decide which open model runs, in which size, and when it is updated.
* **No AI tokens spent.** Requests through your own connection do not use your organisation's Akollo AI tokens. See
  [AI access and tokens](/en/ai/access-and-tokens).
* **It is a good first step.** A laptop is enough to try it with a small model before you plan a server. For a full
  installation inside the institution, see [Running the model in the institution](/en/ai/on-prem).

<Mermaid
  title="From your model server to Akollo features"
  chart="`flowchart LR
A[Your machine or server] --> B[Model runtime]
B --> C[OpenAI-compatible endpoint]
C -->|Directly over HTTPS| D[Own model server connection]
C -->|Through the AI Connector| D
D --> E[Jobs: answers, routing, search]`"
/>

## Choose a runtime [#choose-a-runtime]

Any server that speaks the OpenAI-compatible API works. These three are the most common and are all free to use.

| Runtime                        | Best for                                        | How you run it                             | Embeddings                                            |
| ------------------------------ | ----------------------------------------------- | ------------------------------------------ | ----------------------------------------------------- |
| **llama.cpp** (`llama-server`) | Servers, full control, CPU-only machines        | Command line, one model per server process | Yes, as a separate server started with `--embeddings` |
| **Ollama**                     | Quick setup on a laptop or a small server       | Background service, models pulled by name  | Yes                                                   |
| **LM Studio**                  | Trying models on a desktop with a graphical app | Desktop app with a built-in local server   | Yes                                                   |

For a shared server that the whole organisation uses, llama.cpp (or vLLM on a graphics-card server) is usually the
better choice. Ollama and LM Studio are the easiest way to try things on one machine.

### Install [#install]

<Tabs items="[&#x22;macOS&#x22;,&#x22;Windows&#x22;,&#x22;Linux&#x22;]">
  <Tab value="macOS">
    With Homebrew:

    ```bash
    # llama.cpp (includes llama-server)
    brew install llama.cpp

    # or Ollama
    brew install ollama
    ```

    You can also download the Ollama app from ollama.com. For LM Studio, download the app from lmstudio.ai and drag it
    to Applications. Apple silicon Macs use the built-in graphics memory automatically.
  </Tab>

  <Tab value="Windows">
    With winget, in PowerShell:

    ```powershell
    # llama.cpp (includes llama-server)
    winget install llama.cpp

    # or Ollama
    winget install Ollama.Ollama
    ```

    You can also download the Ollama installer from ollama.com, or a ready-built llama.cpp release from the llama.cpp
    project on GitHub (pick the build that matches your graphics card, or the CPU build). For LM Studio, download the
    installer from lmstudio.ai.
  </Tab>

  <Tab value="Linux">
    Ollama has an official install script:

    ```bash
    curl -fsSL https://ollama.com/install.sh | sh
    ```

    For llama.cpp, download a ready-built release from the llama.cpp project on GitHub, install it with your package
    manager if it offers one (Homebrew on Linux also works), or build it from source:

    ```bash
    git clone https://github.com/ggml-org/llama.cpp
    cd llama.cpp
    cmake -B build
    cmake --build build --config Release
    ```

    Add the build option for your graphics card (for example CUDA) if you have one. LM Studio is available for Linux as
    an AppImage from lmstudio.ai.
  </Tab>
</Tabs>

## Which model for which machine [#which-model-for-which-machine]

The models below are the ones Akollo's own sizing is based on. They are open models you can download from Hugging
Face or pull by name in Ollama and LM Studio.

| Machine                                                         | Model                            | Use it for                                           |
| --------------------------------------------------------------- | -------------------------------- | ---------------------------------------------------- |
| Laptop with 16 GB of memory                                     | **Gemma 4 E4B** or **Jan-v1-4B** | Trying it out, routing, safety checks, short answers |
| Apple silicon with 32–48 GB, or a PC with a 24 GB graphics card | **Gemma 4 26B-A4B**              | Everyday answers and planning for a team or a pilot  |
| Any of the above, alongside the chat model                      | **EmbeddingGemma**               | Search. It is small and runs well on a processor     |

**Quantisation** means storing the model's numbers with fewer bits. A **Q4** (4-bit) file is about half the size of a
**Q8** (8-bit) file, loads faster and answers faster, with a small loss in quality. Q4 is a good default. Pick Q8 only
if the machine has memory to spare. Model files for llama.cpp have the extension `.gguf`, and the quantisation is
usually part of the file name (for example `Q4_K_M`).

<BarChart title="Approximate memory for the model weights at 4-bit" caption="Rough figures for the weights only. Add 30–50% so several people can be served at once, and leave memory for the operating system." unit=" GB" data="[{ label: 'EmbeddingGemma', value: 0.5 }, { label: 'Jan-v1-4B', value: 3 }, { label: 'Gemma 4 E4B', value: 4.5 }, { label: 'Gemma 4 26B-A4B', value: 16 }]" />

For sizing a server for hundreds or thousands of people, see
[Running the model in the institution](/en/ai/on-prem).

## Start the server [#start-the-server]

In the examples, `<port>` is a port you choose and `<model-file>` is the model you downloaded. The OpenAI-compatible
endpoint is the server address followed by `/v1`.

<Tabs items="[&#x22;llama.cpp&#x22;,&#x22;Ollama&#x22;,&#x22;LM Studio&#x22;]">
  <Tab value="llama.cpp">
    Start one server for answers and, if you want to use your own model for search, a second one for embeddings:

    ```bash
    # Chat model
    llama-server -m <model-file>.gguf --port <port>

    # Or let llama-server download a model from Hugging Face
    llama-server -hf <publisher>/<model>-GGUF:Q4_K_M --port <port>

    # Embedding model, on its own port
    llama-server -m <embedding-model-file>.gguf --embeddings --port <port>
    ```

    To protect the server with a key, add `--api-key <your-key>`. Enter the same key in Akollo.
  </Tab>

  <Tab value="Ollama">
    ```bash
    # Start the service (the desktop app starts it for you)
    ollama serve

    # Download a chat model and an embedding model
    ollama pull <model>
    ollama pull <embedding-model>

    # Check what is installed
    ollama list
    ```

    Ollama has no built-in key. If other machines can reach it, put it behind a reverse proxy that checks a key and adds
    HTTPS.
  </Tab>

  <Tab value="LM Studio">
    1. Download a model in the app's model search.
    2. Open **Developer** and start the server.
    3. Load the model you want to serve. The server settings show the address and port.

    You can also start the server from the command line with `lms server start`.
  </Tab>
</Tabs>

### Who can reach the server [#who-can-reach-the-server]

By default, all three runtimes listen only on the machine itself (localhost). That is the safest setting and is
enough when the AI Connector runs on the same machine.

| You want                       | What to do                                                                                                                                                                                      |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Only this machine              | Keep the default.                                                                                                                                                                               |
| Other machines in your network | llama.cpp: add `--host 0.0.0.0`. Ollama: set the `OLLAMA_HOST` environment variable to `0.0.0.0`. LM Studio: turn on serving on the local network in the server settings.                       |
| Akollo connects directly       | Put the server behind HTTPS with a valid certificate (for example a reverse proxy), protect it with a key, and ask your installation's operator to add its address to the allowed destinations. |

<Callout type="warn" title="Warning">
  Never expose a model server to the internet without a key. Anyone who can reach it can use your hardware. With
  llama.cpp use `--api-key`. For Ollama and LM Studio, use a reverse proxy that checks a key, or keep the server
  private and use the AI Connector.
</Callout>

## Connect it in Akollo [#connect-it-in-akollo]

Organisation owners and admins manage connections. See [AI connections](/en/ai/connections) for the whole page.

<Steps>
  <Step>
    ### Open the connections page [#open-the-connections-page]

    Open **Settings › AI › Connections**.
  </Step>

  <Step>
    ### Add the server [#add-the-server]

    Under **Add a connection**, choose **Own model server** as the **Provider** and give it a **Name**. Under **Address**,
    enter the server's HTTPS address including `/v1`, for example `https://<your-model-server>/v1`. If you set a key on
    the server, paste it under **API key**. Otherwise leave it empty.
  </Step>

  <Step>
    ### Choose which data may go through it [#choose-which-data-may-go-through-it]

    Under **Data allowed**, pick **Public data only**, **Up to internal data** or **Up to confidential data**, then choose
    **Add connection**.
  </Step>

  <Step>
    ### Find models and test [#find-models-and-test]

    Choose **Find models**. The list shows the models your server serves. Pick a model and choose **Test**. The test
    sends a short set of fixed requests: answers in Turkish, structured answers and using a tool. The connection can be
    activated only after it passes.
  </Step>

  <Step>
    ### Activate [#activate]

    Choose **Activate**. In live organisations, a second admin approves the activation on the same page.
  </Step>

  <Step>
    ### Choose which jobs use it [#choose-which-jobs-use-it]

    Under **Which model does which job**, choose your connection and model for the jobs you want, then choose **Save**.
    Jobs you leave alone keep using the models Akollo provides.
  </Step>
</Steps>

A sensible split with the models above:

| Job                           | Suggested model                                     |
| ----------------------------- | --------------------------------------------------- |
| **Routing**                   | Jan-v1-4B                                           |
| **Safety check**              | Gemma 4 E4B                                         |
| **Everyday answers**          | Gemma 4 26B-A4B                                     |
| **Planning and complex work** | Gemma 4 26B-A4B, or keep the models Akollo provides |
| **Search**                    | EmbeddingGemma                                      |

<Callout type="info" title="Note">
  A small model may fail the tool or structured-answer checks. Use it for **Routing** or **Safety check** and give
  **Everyday answers** a larger model.
</Callout>

## Servers inside a company network [#servers-inside-a-company-network]

From a cloud installation, Akollo does not connect to private network addresses. If your model server sits inside
your company network, or you do not want to open an inbound port, use the **AI Connector**. It is a small container
that runs inside your network next to the model server. It connects out to Akollo over HTTPS, picks up requests, sends
them to the model server and returns the answers. No inbound port is opened.

<Mermaid
  title="AI Connector inside a company network"
  chart="`flowchart LR
subgraph Company network
M[Model server] --- K[AI Connector]
end
K -->|Outbound HTTPS with its token| A[Akollo]`"
/>

<Steps>
  <Step>
    ### Add the connector in Akollo [#add-the-connector-in-akollo]

    On **Settings › AI › Connections**, under **Add an AI Connector**, give it a name and choose **Add connector**.
  </Step>

  <Step>
    ### Copy the token [#copy-the-token]

    Copy the **Connector token**. It is shown only once. Keep it in your secret store.
  </Step>

  <Step>
    ### Start the container [#start-the-container]

    Run the AI Connector container next to the model server. It needs three settings: Akollo's address, the token, and
    the model server's OpenAI-compatible address. A fourth, optional setting holds the model server's key if it needs one.

    ```bash
    docker run --restart unless-stopped \
      -e AKOLLO_URL=https://<your-akollo-address> \
      -e AKOLLO_CONNECTOR_TOKEN=<connector-token> \
      -e AKOLLO_CONNECTOR_TARGET=http://<model-server>:<port>/v1 \
      -e AKOLLO_CONNECTOR_TARGET_KEY=<model-server-key> \
      <ai-connector-image>
    ```
  </Step>

  <Step>
    ### Check, test and activate [#check-test-and-activate]

    The connection card shows **Connector online** once it connects. Then use **Find models**, **Test** and **Activate**
    as for a direct server.
  </Step>
</Steps>

* The connector only needs outgoing HTTPS to Akollo and a route to the model server. The model server itself can stay
  on plain HTTP inside your network.
* Run one copy per token, restarted automatically if it stops. A second copy with the same token is refused.
* The connector forwards model listing and chat answers. For the **Search** job, use a direct **Own model server**
  connection or the models Akollo provides.
* **New token** stops the old token at once. Restart the connector with the new token, then test and activate the
  connection again.
* The connector logs one line per request and never the token or the content.

## Troubleshooting [#troubleshooting]

| Symptom                                           | Likely cause                                                                    | Fix                                                                                                                                                           |
| ------------------------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The test says the connection could not be reached | Wrong address, missing `/v1`, wrong key, or the server is not running           | Open the address followed by `/models` from a machine that should reach it. Check the key, then test again.                                                   |
| The address is not accepted                       | Not HTTPS, not on the installation's allowed list, or a private network address | Put the server behind HTTPS and ask your operator to allow the address, or use the AI Connector.                                                              |
| **Find models** shows no models                   | No model loaded, or the key belongs to another server                           | Load or pull a model (in LM Studio, load it in the server view), then choose **Find models** again.                                                           |
| The test fails some checks                        | The model is too small for structured answers or tool use                       | Use a larger model, such as Gemma 4 26B-A4B, for **Everyday answers**. Keep small models for **Routing** and **Safety check**.                                |
| Answers are slow or time out                      | The model is too large for the machine, or it runs only on the processor        | Use a smaller model or a Q4 file, make sure the graphics card is used, and close other heavy programs.                                                        |
| The server stops with an out-of-memory error      | Model plus context does not fit in memory                                       | Use a smaller model or quantisation, or shorten the context window.                                                                                           |
| Search results are poor or fail                   | A chat model was chosen for **Search**, or the embedding server is not running  | Choose an embedding model such as EmbeddingGemma for **Search**. With llama.cpp, start it with `--embeddings`. Do not switch embedding models back and forth. |
| Certificate or TLS errors                         | Self-signed or expired certificate                                              | Use a certificate your installation trusts, or use the AI Connector so Akollo does not connect to the server directly.                                        |
| Browser or CORS errors when testing by hand       | The runtime blocks requests from web pages                                      | This does not affect Akollo, which calls the server from its own servers. Test with a command-line tool instead.                                              |
| Other machines cannot reach the server            | It listens only on localhost, or a firewall blocks the port                     | Bind it to the network (see above) and open the port only to the machines that need it.                                                                       |
| **Connector offline**                             | Wrong token, Akollo's address not reachable, or the token was replaced          | Check the container log, the three settings and outgoing HTTPS. After **New token**, restart with the new token.                                              |

## FAQ [#faq]

<Accordions type="single">
  <Accordion title="Do I need a graphics card?">
    No. Small models such as Jan-v1-4B, Gemma 4 E4B and EmbeddingGemma run on a processor. A graphics card, or an Apple
    silicon Mac with enough memory, makes the larger Gemma 4 26B-A4B fast enough for everyday answers.
  </Accordion>

  <Accordion title="Do requests to my own model use our AI tokens?">
    No. Requests through your own connection do not use your organisation's Akollo AI tokens.
  </Accordion>

  <Accordion title="Can I try it on my laptop first?">
    Yes. Install Ollama or LM Studio, load a small model, and connect it through the AI Connector running on the same
    laptop. Use a model your laptop can hold, such as Gemma 4 E4B.
  </Accordion>

  <Accordion title="Do I have to open a port in our firewall?">
    Not with the AI Connector. It connects out to Akollo, so no inbound port is opened. A direct **Own model server**
    connection needs an HTTPS address that Akollo can reach.
  </Accordion>

  <Accordion title="Can I use my own model for some jobs and Akollo's models for others?">
    Yes. Under **Which model does which job**, choose a model for each job separately. Jobs you leave alone keep using
    the models Akollo provides.
  </Accordion>

  <Accordion title="What happens if my server goes down?">
    Requests that use it fail until it is back. To move jobs back to the models Akollo provides, change them under
    **Which model does which job**. **Disable** turns the connection off for good.
  </Accordion>
</Accordions>

## Related pages [#related-pages]

* [AI connections](/en/ai/connections)
* [Models that stay in the institution](/en/ai/local-models)
* [Running the model in the institution](/en/ai/on-prem)
* [What assistance can see](/en/ai/privacy)
* [AI access and tokens](/en/ai/access-and-tokens)
