How to connect Cherry Studio with Ollama step by step

Last update: 26/08/2026

  • Cherry Studio expands Ollama by adding multiple providers, advanced assistants, and RAG knowledge bases.
  • The connection requires enabling the native provider in Cherry Studio using the local URL http://localhost:11434.
  • You must register the exact name of the downloaded model, respecting its labels and colon.
  • Optimal local performance requires choosing appropriate models based on the available VRAM and RAM memory.
How to connect Cherry Studio with Ollama

Ollama It greatly simplifies the download and execution of artificial intelligence models on a computer. You can start a conversation from the terminal, use its official application, or connect other programs via the local API that runs in the background. However, when you want to organize many conversations, switch between providers, create specialized assistants, or work with knowledge bases, An application like Cherry Studio can offer a more complete environment.

In this guide we will see How to connect Cherry Studio with Ollamawhich address to use, how to add downloaded models, and what to do if the provider appears offline or performance is too low.

What does Cherry Studio offer if Ollama already has an app?

What Cherry Studio offers

Ollama is no longer just a terminal tool. Its Windows and macOS applications allow you to download models, hold conversations, and attach files. Therefore, you don't necessarily need to install a separate interface to start using local AI.

Cherry Studio is a great option when you're looking for a more comprehensive workspace. The application allows you to bring together local models and cloud providers, organize attendees, manage conversations, prepare reusable instructions, and work with knowledge bases.

The difference can be summarized as follows:

Function Ollama application Cherry Studio
Download and run models Yeah Use the models managed by Ollama
Basic conversations Yeah Yeah
Attach files Yeah Yeah
Combining local models and APIs It also supports Ollama cloud models It allows you to configure numerous different providers.
Personalized assistants Easier setup Greater organization and personalization
Knowledge bases That is not its primary function Yes, through documents, embeddings, and retrieval

Cherry Studio Community Edition is free and distributed under the AGPL-3.0 license. This does not mean that all models available within the application are free: APIs from OpenAI, Anthropic, Google, or other providers may charge for their use.

Requirements to connect Cherry Studio with Ollama

Before opening Cherry Studio, check that Ollama is installed and that your local server is responding correctly.

The easiest way to review the available models is:

ollama list

You can also directly check the Ollama endpoint:

curl http://localhost:11434/api/tags

In PowerShell, it may be more reliable to explicitly specify the executable:

curl.exe http://localhost:11434/api/tags

If Ollama is working, you will receive a JSON object with the installed models. You can also open this address in your browser:

http://localhost:11434

When the service is available, Ollama displays a message indicating that it is running.

If the connection fails, start or restart Ollama depending on your system:

  • Windows: Open Ollama from the Start menu or restart it from the tray icon.
  • macOS: Close and reopen the Ollama application.
  • Linux with systemd: executes sudo systemctl restart ollama.
  • Manual execution: Stop the active process and run it again. ollama serve.

It's also recommended to use a recent version. Ollama receives frequent updates to add models, fix bugs, and expand API compatibility. However, Cherry Studio doesn't specifically require Ollama 0.13.3 when using its native provider.

How to choose a suitable local model

How to choose a suitable local model

The model should be chosen according to the RAM, the VRAM of the graphics card, the context you want to use it in, and the speed you consider acceptable.

Exclusive content - Click Here  How to use Stable Diffusion 3 on your PC: requirements and recommended models

For a first test, qwen3:4b It's a relatively lightweight option:

ollama pull qwen3:4b

The file published in the Ollama library is approximately 2,5 GB in size. Actual memory consumption during execution will be higher because memory must also be reserved for context, cache, and other model data.

If you have more memory, you can try:

ollama pull qwen3:8b

This variant takes up approximately 5,2 GB, although it will require additional memory when running.

For a machine with around 24 GB of VRAM, it may be possible to use larger models, such as:

ollama pull qwen3:30b

This variant occupies approximately 19 GB. Always leave room for the context cache and other applications. A model that fits as a file in VRAM can run out of memory when you increase the context or run multiple conversations simultaneously.

A purported version of Call 4 of 8B Because it doesn't exist. Llama 4 Scout and Maverick are considerably larger MoE models. The version of Scout released on Llama takes up around 67 GB, so it's not a practical option for most home computers.

Before setting up Cherry Studio, check that the model responds correctly:

ollama run qwen3:4b

Write a short question and use /bye to end the conversation.

How to use several different models in Ollama without conflicts
Related article:
Complete Guide to Managing Multiple Models in Ollama Without Conflicts

How to connect Cherry Studio with Ollama step by step

Ollama provider configuration within Cherry Studio

Download Cherry Studio from its official website or the project's GitHub repository. Install the version corresponding to Windows, macOS, or Linux and open the application.

  1. Enter Settings.
  2. Open the section Model providers o Modeling services, depending on the version.
  3. Locate the supplier Ollama and activate it.
  4. Enter http://localhost:11434 as an API address.
  5. Leave the API key empty if the interface allows it.
  6. Open the model management.
  7. Add the exact name that appears when you run ollama list.
  8. Save the settings and run a test.

For example, if you have downloaded:

ollama pull qwen3:4b

The identifier you need to add in Cherry Studio will be:

qwen3:4b

Respect the colon and the etiquette. Write only qwen3 You can select another variant or prevent Cherry Studio from finding the expected model.

When should you add /v1 and a dummy API key?

The above configuration uses Ollama's native provider. In that case, the recommended address is:

http://localhost:11434

If Ollama is not listed or you want to configure it as a custom provider compatible with OpenAI, the values ​​change:

URL base: http://localhost:11434/v1
API key: ollama

The key ollama It is not checked. Some OpenAI-compatible libraries require the field to contain a value, even though Ollama does not use that key for local connections.

Do not mix the two settings. Add /v1 Using the native provider may cause errors if Cherry Studio expects to use Ollama's own endpoints.

How to use files and knowledge bases

Cherry Studio allows you to attach documents to a conversation and create knowledge bases to retrieve snippets related to each question.

A knowledge base is not simply about dragging and dropping a PDF. The process typically requires:

  • The original documents.
  • A system for dividing the text into fragments.
  • An embeddings model.
  • A vector warehouse.
  • Optionally, a reranking model.
  • The generative model that drafts the response.
Exclusive content - Click Here  How to install plugins in Dify

If you want the process to remain on your computer, you must select local models for both generation and embedding, as well as any additional functions. Using Ollama to respond, but an external API to generate embeds, means that some content will be sent to the corresponding provider.

You should also check if you have backups, WebDAV synchronization, web search, MCP, or other integrations enabled that may communicate with external services.

RAG can help substantiate the answers in your documents, but it doesn't guarantee that all results will be correct. Review the retrieved sources and don't accept an answer as valid simply because a knowledge base was used.

How to fix connection errors

How to fix Cherry Studio connection errors with Ollama

Cherry Studio shows Ollama as offline

First, check these addresses:

http://localhost:11434
http://localhost:11434/api/tags

If Ollama is working, check that Cherry Studio is using the correct URL for the selected provider type.

  • Ollama Supplier: http://localhost:11434
  • OpenAI compatible provider: http://localhost:11434/v1

The model does not appear in Cherry Studio

Execute:

ollama list

Manually add the full identifier that appears in the column corresponding to the name in Cherry Studio.

The error "model not found" appears.

This usually means that Cherry Studio is requesting an identifier that Ollama doesn't have installed. Download it or correct the name:

ollama pull qwen3:4b

The connection returns a 404 error

Check if you have added or removed /v1 incorrectly. This can also happen when a Cherry Studio function uses an endpoint that the installed version of Ollama doesn't yet support. Update both programs and try again.

Port 11434 is occupied

Do not start multiple instances of ollama serve simultaneously. If the Ollama application is already active, you don't need to open another server manually.

How to improve Ollama's performance

If the model is taking too long, first check where it is loaded:

ollama ps

The column PROCESSOR It may display results such as:

  • 100% GPUThe model is fully loaded onto the GPU.
  • 100% CPU: is running in system memory.
  • CPU/GPUPart of it is in the GPU and part is in RAM.

On systems with NVIDIA you can also use:

nvidia-smi

This command allows you to see the process, GPU usage, and memory used, but ollama ps It shows more clearly the distribution used by Ollama.

If the model doesn't fit entirely in VRAM, some of the load may be shifted to RAM or the CPU. This significantly reduces speed. Try using a smaller model or decreasing the context length.

You can also close memory-intensive applications, especially browsers with numerous tabs, virtual machines, video editors, and other programs running simultaneously.

How to prevent the model from being downloaded from memory

Ollama keeps the models loaded for a certain period after each request. You can modify this using OLLAMA_KEEP_ALIVE.

For example, to keep them for thirty minutes:

OLLAMA_KEEP_ALIVE=30m

You can also use:

OLLAMA_KEEP_ALIVE=-1

A negative value keeps the model loaded indefinitely. This reduces the wait time for the next query, but also keeps RAM or VRAM occupiedIt is not recommended if you use multiple models, play games, edit video, or need to free up memory for other applications.

The parameter keep_alive The command sent in a specific request can override the global variable. Some versions of Cherry Studio also allow you to adjust this value from the provider settings.

Cherry Studio versus other interfaces for Ollama

Cherry Studio isn't the only application capable of using models managed by Ollama. Other available alternatives include Levante, Msty, LM Studio, Open WebUI, and Ollama's own application.

Exclusive content - Click Here  Complete Guide to Reviewing Contracts with Artificial Intelligence
Application Strength Considerations
Cherry Studio Numerous suppliers, assistants, and knowledge bases Some features require additional configuration.
Ollama Easy installation and direct engine management Less geared towards a multi-vendor workspace
Levant Local environment, multiple providers, and MCP Your license should be reviewed according to the intended use
Msty Organizing conversations, understanding and comparing models Combine free features with other license terms or plans
LM Studio Download, configure, and run models from a single application It can replace Ollama as an engine instead of just using it
Open WebUI Web interface suitable for servers and various devices Its deployment usually requires Docker or another server installation

Cherry Studio is particularly suitable if you want to switch between Ollama and different APIs from the same application. If you only need a simple local conversation, the official Ollama application may suffice.

Mastering LM Studio
Related article:
Complete Guide to LM Studio: How to Master Local AI on Your PC

How to use Ollama from another device on the network

How to use Ollama from another device on the network

Ollama listens by default on 127.0.0.1:11434Therefore, it only accepts connections from the same device.

To make it accessible from the local network, configure:

OLLAMA_HOST=0.0.0.0:11434

The procedure depends on the operating system.

Windows

  1. Close Ollama from the system tray.
  2. Open the environment variables settings.
  3. Create a user variable called OLLAMA_HOST.
  4. Enter 0.0.0.0:11434 as a value.
  5. Start Ollama again.

Linux with systemd

Execute:

sudo systemctl edit ollama.service

Add:

[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"

Then apply the change:

sudo systemctl daemon-reload
sudo systemctl restart ollama

macOS

Configure the variable and restart the application:

launchctl setenv OLLAMA_HOST "0.0.0.0:11434"

In Cherry Studio, replace localhost by the IP address of the computer running Ollama:

http://192.168.1.50:11434

Ollama does not require authentication for this local endpoint. Do not open port 11434 to the Internet.

Configure your firewall to accept connections only from the necessary devices or subnet. If you want to access the site from outside your home network, use a private VPN or a reverse proxy with authentication and HTTPS.

How to create a custom model using a Modelfile

Un Modelfile It allows you to create a variant with predefined instructions and parameters.

Create a file called Modelfile with similar content:

FROM qwen3:4b

SYSTEM """
Responde en español.
Explica los pasos con claridad.
No inventes comandos ni datos técnicos.
"""

PARAMETER temperature 0.3

Then create the model:

ollama create asistente-tecnico -f Modelfile

Check that it appears on the list:

ollama list

Finally, he adds asistente-tecnico in Cherry Studio's model management. This way, you'll have the default instructions available every time you select that variant.

Connecting Cherry Studio with Ollama allows you to bring together conversations, attendees, documents, and different providers within a single application. The simplest setup involves activating the native provider and using http://localhost:11434 and add the exact identifier of the downloaded model.

To maintain a truly local environment, remember to use local models at all stages, including embedding and reranking. You should also avoid external synchronizations and integrations if the documents contain sensitive information.

If performance is low, use ollama psReduce the model size or decrease the context. And if you're going to share Ollama on the network, protect the port so that no unauthorized device can consume your computer's resources.