Software:Ollama

From HandWiki
Ollama
Developer(s)Ollama Inc. and contributors
Initial releaseJuly 7, 2023; 3 years ago (2023-07-07)
Stable release
0.22.1 / April 28, 2026; 4 months ago (2026-04-28)
Repositorygithub.com/ollama/ollama
Written inGo
Operating systemmacOS, Linux, Windows
TypeLarge language model runtime
LicenseMIT License
Websiteollama.com

Ollama is an open-source software platform developed by Jeffrey Morgan and Michael Chiang in 2023 for running and managing large language models on local computers and through hosted cloud models. It provides a command-line interface, a native GUI, a local REST API, model-management tools, and integrations for using open-weight models with coding assistants and other applications.[1][2][3]

History

Ollama was first released in 2023.[4] The project became associated with the growth of local large language model software, allowing users to download and run models such as Llama, Gemma, Mistral, Qwen, gpt-oss, GLM and DeepSeek on a local machine.[1]

In 2025 and 2026, Ollama added additional application and cloud features, including hosted cloud models, web search support, tool and coding-agent integrations, and support for using Ollama with applications such as Claude Code, Codex, OpenCode, Copilot CLI,[5] and OpenClaw.[2][6] In March 2026, Ollama announced preview support for Apple silicon.[7]

In July 2026 it raised $65 million dollars in funding.[8]

Features

Ollama running Llama 3 in Linux

Ollama includes tools for downloading, running, importing, and managing large language models. Users can run models from the command line, interact with them through a local HTTP API, or use client libraries for programming languages such as Python and JavaScript.[1]

The project provides a REST API for chat and model-management functions, with the default local service commonly exposed on port 11434.[9] Ollama also distributes an official Docker image and provides model libraries and documentation for running supported models.[1]

Ollama uses the llama.cpp backend for local model inference,[10] which supports running of quantized models that use less memory. It can employ graphics cards (GPUs) to speed up calculations.[11] It supports a model-library format that allows users to pull, run, and manage model variants by name.[12]

Meanwhile, the project's downloadable software is being explored as a back-end platform for low-code/no-code self-hosting of LLMs, potentially allowing natural language processing use cases to become more accessible to nontechnical end users.[13][14]

Ollama also has experimental support for image generation, using models such as Flux, on macOS, with Windows and Linux support expected in the future.[15]

Security

Because Ollama is commonly used to run local or self-hosted AI models, security researchers have examined risks from misconfigured public deployments. In January 2026, The Hacker News reported on research by SentinelOne and Censys that found many Ollama servers were exposed to the public internet, mainly by binding to the "any" port 0.0.0.0 which on an inadequately-secured system would allow access from other local or remote devices despite the fact that Ollama is intended to run locally by default.[16]

See also

References

  1. ↑ 1.0 1.1 1.2 1.3 "ollama/ollama". https://github.com/ollama/ollama. 
  2. ↑ 2.0 2.1 "Ollama". https://ollama.com/. 
  3. ↑ SitePoint Team (April 22, 2026). "10GB VRAM Local LLM: The Complete Setup Guide (2026)". https://www.sitepoint.com/10gb-vram-local-llm-the-complete-setup-guide-2026/. 
  4. ↑ "Versions: github.com/ollama/ollama". https://deps.dev/go/github.com%2Follama%2Follama/v0.1.29/versions. 
  5. ↑ "About GitHub Copilot CLI". GitHub. https://docs.github.com/en/copilot/concepts/agents/copilot-cli/about-copilot-cli. 
  6. ↑ "Blog". https://ollama.com/blog. 
  7. ↑ Medina, David (March 31, 2026). "Ollama Taps Apple's MLX Framework to Make Local AI Faster". https://thenewstack.io/ollama-taps-apples-mlx/. 
  8. ↑ Bort, Julie (2026-07-09). "Popular open source AI developer tool Ollama raises $65M, grows to nearly 9M users" (in en-US). https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users/. 
  9. ↑ "API reference". https://docs.ollama.com/api. 
  10. ↑ "Forensic Implications of Localized AI: Artifact Analysis of Ollama, LM Studio, and llama.cpp". https://arxiv.org/html/2603.23996v1. 
  11. ↑ Easton (2026-04-25). "Ollama GPU Setup: CUDA, ROCm, Metal, and Troubleshooting" (in en). https://eastondev.com/blog/en/posts/ai/20260425-ollama-gpu-acceleration/. 
  12. ↑ "Library". https://ollama.com/library. 
  13. ↑ Sempio, Julius (2026-03-31). "Low-Code Self-Hosting of RAG-Enabled Language Models Using Ollama and Open WebUI" (in en). doi:10.1109/ICAIIC68212.2026.11454259. https://ieeexplore.ieee.org/abstract/document/11454259. 
  14. ↑ "Low-Code vs. No-Code: What’s the difference?". https://www.ibm.com/think/topics/low-code-vs-no-code. 
  15. ↑ Kemper, Jonathan (2026-01-21). "Ollama brings local AI image generation to Mac" (in en-US). https://the-decoder.com/ollama-brings-local-ai-image-generation-to-mac/. 
  16. ↑ Lakshmanan, Ravie (January 29, 2026). "Researchers Find 175,000 Publicly Exposed Ollama AI Servers Across 130 Countries". https://thehackernews.com/2026/01/researchers-find-175000-publicly.html. 

Template:Large language models