- Cherry Studio expands Ollama by adding multiple providers, advanced assistants, and RAG knowledge bases.
- The connection requires enabling the native provider in Cherry Studio using the local URL http://localhost:11434.
- You must register the exact name of the downloaded model, respecting its labels and colon.
- Optimal local performance requires choosing appropriate models based on the available VRAM and RAM memory.
Ollama It greatly simplifies the download and execution of artificial intelligence models on a computer. You can start a conversation from the terminal, use its official application, or connect other programs via the local API that runs in the background. However, when you want to organize many conversations, switch between providers, create specialized assistants, or work with knowledge bases, An application like Cherry Studio can offer a more complete environment.
In this guide we will see How to connect Cherry Studio with Ollamawhich address to use, how to add downloaded models, and what to do if the provider appears offline or performance is too low.
What does Cherry Studio offer if Ollama already has an app?

Ollama is no longer just a terminal tool. Its Windows and macOS applications allow you to download models, hold conversations, and attach files. Therefore, you don't necessarily need to install a separate interface to start using local AI.
Cherry Studio is a great option when you're looking for a more comprehensive workspace. The application allows you to bring together local models and cloud providers, organize attendees, manage conversations, prepare reusable instructions, and work with knowledge bases.
The difference can be summarized as follows:
| Function | Ollama application | Cherry Studio |
|---|---|---|
| Download and run models | Yeah | Use the models managed by Ollama |
| Basic conversations | Yeah | Yeah |
| Attach files | Yeah | Yeah |
| Combining local models and APIs | It also supports Ollama cloud models | It allows you to configure numerous different providers. |
| Personalized assistants | Easier setup | Greater organization and personalization |
| Knowledge bases | That is not its primary function | Yes, through documents, embeddings, and retrieval |
Cherry Studio Community Edition is free and distributed under the AGPL-3.0 license. This does not mean that all models available within the application are free: APIs from OpenAI, Anthropic, Google, or other providers may charge for their use.
Requirements to connect Cherry Studio with Ollama
Before opening Cherry Studio, check that Ollama is installed and that your local server is responding correctly.
The easiest way to review the available models is:
ollama list
You can also directly check the Ollama endpoint:
curl http://localhost:11434/api/tags
In PowerShell, it may be more reliable to explicitly specify the executable:
curl.exe http://localhost:11434/api/tags
If Ollama is working, you will receive a JSON object with the installed models. You can also open this address in your browser:
http://localhost:11434
When the service is available, Ollama displays a message indicating that it is running.
If the connection fails, start or restart Ollama depending on your system:
- Windows: Open Ollama from the Start menu or restart it from the tray icon.
- macOS: Close and reopen the Ollama application.
- Linux with systemd: executes
sudo systemctl restart ollama. - Manual execution: Stop the active process and run it again.
ollama serve.
It's also recommended to use a recent version. Ollama receives frequent updates to add models, fix bugs, and expand API compatibility. However, Cherry Studio doesn't specifically require Ollama 0.13.3 when using its native provider.
How to choose a suitable local model

The model should be chosen according to the RAM, the VRAM of the graphics card, the context you want to use it in, and the speed you consider acceptable.
For a first test, qwen3:4b It's a relatively lightweight option:
ollama pull qwen3:4b
The file published in the Ollama library is approximately 2,5 GB in size. Actual memory consumption during execution will be higher because memory must also be reserved for context, cache, and other model data.
If you have more memory, you can try:
ollama pull qwen3:8b
This variant takes up approximately 5,2 GB, although it will require additional memory when running.
For a machine with around 24 GB of VRAM, it may be possible to use larger models, such as:
ollama pull qwen3:30b
This variant occupies approximately 19 GB. Always leave room for the context cache and other applications. A model that fits as a file in VRAM can run out of memory when you increase the context or run multiple conversations simultaneously.
A purported version of Call 4 of 8B Because it doesn't exist. Llama 4 Scout and Maverick are considerably larger MoE models. The version of Scout released on Llama takes up around 67 GB, so it's not a practical option for most home computers.
Before setting up Cherry Studio, check that the model responds correctly:
ollama run qwen3:4b
Write a short question and use /bye to end the conversation.
How to connect Cherry Studio with Ollama step by step

Download Cherry Studio from its official website or the project's GitHub repository. Install the version corresponding to Windows, macOS, or Linux and open the application.
- Enter Settings.
- Open the section Model providers o Modeling services, depending on the version.
- Locate the supplier Ollama and activate it.
- Enter
http://localhost:11434as an API address. - Leave the API key empty if the interface allows it.
- Open the model management.
- Add the exact name that appears when you run
ollama list. - Save the settings and run a test.
For example, if you have downloaded:
ollama pull qwen3:4b
The identifier you need to add in Cherry Studio will be:
qwen3:4b
Respect the colon and the etiquette. Write only qwen3 You can select another variant or prevent Cherry Studio from finding the expected model.
When should you add /v1 and a dummy API key?
The above configuration uses Ollama's native provider. In that case, the recommended address is:
http://localhost:11434
If Ollama is not listed or you want to configure it as a custom provider compatible with OpenAI, the values change:
URL base: http://localhost:11434/v1
API key: ollama
The key ollama It is not checked. Some OpenAI-compatible libraries require the field to contain a value, even though Ollama does not use that key for local connections.
Do not mix the two settings. Add /v1 Using the native provider may cause errors if Cherry Studio expects to use Ollama's own endpoints.
How to use files and knowledge bases
Cherry Studio allows you to attach documents to a conversation and create knowledge bases to retrieve snippets related to each question.
A knowledge base is not simply about dragging and dropping a PDF. The process typically requires:
- The original documents.
- A system for dividing the text into fragments.
- An embeddings model.
- A vector warehouse.
- Optionally, a reranking model.
- The generative model that drafts the response.
If you want the process to remain on your computer, you must select local models for both generation and embedding, as well as any additional functions. Using Ollama to respond, but an external API to generate embeds, means that some content will be sent to the corresponding provider.
You should also check if you have backups, WebDAV synchronization, web search, MCP, or other integrations enabled that may communicate with external services.
RAG can help substantiate the answers in your documents, but it doesn't guarantee that all results will be correct. Review the retrieved sources and don't accept an answer as valid simply because a knowledge base was used.
How to fix connection errors

Cherry Studio shows Ollama as offline
First, check these addresses:
http://localhost:11434
http://localhost:11434/api/tags
If Ollama is working, check that Cherry Studio is using the correct URL for the selected provider type.
- Ollama Supplier:
http://localhost:11434 - OpenAI compatible provider:
http://localhost:11434/v1
The model does not appear in Cherry Studio
Execute:
ollama list
Manually add the full identifier that appears in the column corresponding to the name in Cherry Studio.
The error "model not found" appears.
This usually means that Cherry Studio is requesting an identifier that Ollama doesn't have installed. Download it or correct the name:
ollama pull qwen3:4b
The connection returns a 404 error
Check if you have added or removed /v1 incorrectly. This can also happen when a Cherry Studio function uses an endpoint that the installed version of Ollama doesn't yet support. Update both programs and try again.
Port 11434 is occupied
Do not start multiple instances of ollama serve simultaneously. If the Ollama application is already active, you don't need to open another server manually.
How to improve Ollama's performance
If the model is taking too long, first check where it is loaded:
ollama ps
The column PROCESSOR It may display results such as:
100% GPUThe model is fully loaded onto the GPU.100% CPU: is running in system memory.CPU/GPUPart of it is in the GPU and part is in RAM.
On systems with NVIDIA you can also use:
nvidia-smi
This command allows you to see the process, GPU usage, and memory used, but ollama ps It shows more clearly the distribution used by Ollama.
If the model doesn't fit entirely in VRAM, some of the load may be shifted to RAM or the CPU. This significantly reduces speed. Try using a smaller model or decreasing the context length.
You can also close memory-intensive applications, especially browsers with numerous tabs, virtual machines, video editors, and other programs running simultaneously.
How to prevent the model from being downloaded from memory
Ollama keeps the models loaded for a certain period after each request. You can modify this using OLLAMA_KEEP_ALIVE.
For example, to keep them for thirty minutes:
OLLAMA_KEEP_ALIVE=30m
You can also use:
OLLAMA_KEEP_ALIVE=-1
A negative value keeps the model loaded indefinitely. This reduces the wait time for the next query, but also keeps RAM or VRAM occupiedIt is not recommended if you use multiple models, play games, edit video, or need to free up memory for other applications.
The parameter keep_alive The command sent in a specific request can override the global variable. Some versions of Cherry Studio also allow you to adjust this value from the provider settings.
Cherry Studio versus other interfaces for Ollama
Cherry Studio isn't the only application capable of using models managed by Ollama. Other available alternatives include Levante, Msty, LM Studio, Open WebUI, and Ollama's own application.
| Application | Strength | Considerations |
|---|---|---|
| Cherry Studio | Numerous suppliers, assistants, and knowledge bases | Some features require additional configuration. |
| Ollama | Easy installation and direct engine management | Less geared towards a multi-vendor workspace |
| Levant | Local environment, multiple providers, and MCP | Your license should be reviewed according to the intended use |
| Msty | Organizing conversations, understanding and comparing models | Combine free features with other license terms or plans |
| LM Studio | Download, configure, and run models from a single application | It can replace Ollama as an engine instead of just using it |
| Open WebUI | Web interface suitable for servers and various devices | Its deployment usually requires Docker or another server installation |
Cherry Studio is particularly suitable if you want to switch between Ollama and different APIs from the same application. If you only need a simple local conversation, the official Ollama application may suffice.
How to use Ollama from another device on the network

Ollama listens by default on 127.0.0.1:11434Therefore, it only accepts connections from the same device.
To make it accessible from the local network, configure:
OLLAMA_HOST=0.0.0.0:11434
The procedure depends on the operating system.
Windows
- Close Ollama from the system tray.
- Open the environment variables settings.
- Create a user variable called
OLLAMA_HOST. - Enter
0.0.0.0:11434as a value. - Start Ollama again.
Linux with systemd
Execute:
sudo systemctl edit ollama.service
Add:
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Then apply the change:
sudo systemctl daemon-reload
sudo systemctl restart ollama
macOS
Configure the variable and restart the application:
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"
In Cherry Studio, replace localhost by the IP address of the computer running Ollama:
http://192.168.1.50:11434
Ollama does not require authentication for this local endpoint. Do not open port 11434 to the Internet.
Configure your firewall to accept connections only from the necessary devices or subnet. If you want to access the site from outside your home network, use a private VPN or a reverse proxy with authentication and HTTPS.
How to create a custom model using a Modelfile
Un Modelfile It allows you to create a variant with predefined instructions and parameters.
Create a file called Modelfile with similar content:
FROM qwen3:4b
SYSTEM """
Responde en español.
Explica los pasos con claridad.
No inventes comandos ni datos técnicos.
"""
PARAMETER temperature 0.3
Then create the model:
ollama create asistente-tecnico -f Modelfile
Check that it appears on the list:
ollama list
Finally, he adds asistente-tecnico in Cherry Studio's model management. This way, you'll have the default instructions available every time you select that variant.
Connecting Cherry Studio with Ollama allows you to bring together conversations, attendees, documents, and different providers within a single application. The simplest setup involves activating the native provider and using http://localhost:11434 and add the exact identifier of the downloaded model.
To maintain a truly local environment, remember to use local models at all stages, including embedding and reranking. You should also avoid external synchronizations and integrations if the documents contain sensitive information.
If performance is low, use ollama psReduce the model size or decrease the context. And if you're going to share Ollama on the network, protect the port so that no unauthorized device can consume your computer's resources.
I am a technology enthusiast who has turned his "geek" interests into a profession. I have spent more than 10 years of my life using cutting-edge technology and tinkering with all kinds of programs out of pure curiosity. Now I have specialized in computer technology and video games. This is because for more than 5 years I have been writing for various websites on technology and video games, creating articles that seek to give you the information you need in a language that is understandable to everyone.
If you have any questions, my knowledge ranges from everything related to the Windows operating system as well as Android for mobile phones. And my commitment is to you, I am always willing to spend a few minutes and help you resolve any questions you may have in this internet world.