🔧 Lesson: Pulling Models (e.g., ollama pull mistral)
🎯 Lesson Objective
By the end of this lesson, you’ll understand how to:
-
Download and manage open-source LLMs locally using Ollama
-
Pull specific models like Mistral, LLaMA, or Code LLaMA
-
Choose models based on your use case (e.g., chat, coding, summarization)
🧠 Why Pull a Model?
Before using a local LLM with Ollama, you need to download (or “pull”) it.
Pulling a model means you’re downloading a pre-trained large language model (LLM) to your machine so it can run without the internet.
Running models locally allows:
-
💡 Full offline capability
-
🕵️♂️ Private document analysis
-
💸 No API costs (unlike OpenAI)
-
⚡ Fast responses (if your hardware supports it)
🚀 Step-by-Step: Pulling a Model Using Ollama
✅ 1. Install Ollama
If you haven’t yet, install Ollama from https://ollama.com.
Supported OS: macOS, Windows, Linux
Installation is a one-line command (platform-specific)
✅ 2. Open Your Terminal or Command Prompt
On your machine, open:
-
macOS/Linux: Terminal
-
Windows: Command Prompt or PowerShell
✅ 3. Pull a Model
Use the command:
ollama pull mistral
This downloads the Mistral 7B model — a powerful, lightweight LLM suitable for:
-
Text generation
-
Summarization
-
Question answering
-
Document chat
-
General-purpose agent building
You’ll see output like:
✔ downloading model mistral:latest from ollama...
✔ success
✅ 4. Run the Model
Once pulled, you can run it with:
ollama run mistral
This starts a chat session with the Mistral model.
🧰 Popular Models You Can Pull
| Model | Use Case | Command |
|---|---|---|
mistral |
General-purpose, fast, accurate | ollama pull mistral |
llama3 |
Meta’s LLaMA 3 model | ollama pull llama3 |
gemma |
Google’s lightweight model | ollama pull gemma |
codellama |
Code-focused tasks | ollama pull codellama |
deepseek-coder |
Strong code generation | ollama pull deepseek-coder |
orca-mini |
Small + fast (low RAM) | ollama pull orca-mini |
⚠️ Model Size & RAM
| Model | Size | RAM Required (approx.) |
|---|---|---|
| Mistral | ~4–8 GB | 8–16 GB |
| LLaMA 3 | ~8–12 GB | 16+ GB |
| Code LLaMA | 8–12 GB | 16+ GB |
| Orca Mini | 3–5 GB | 4–8 GB |
If your system has low RAM, prefer smaller models like:
ollama pull orca-mini
🧪 Troubleshooting
| Issue | Solution |
|---|---|
| “Ollama not found” | Reinstall Ollama and ensure it’s added to your PATH |
| “Model pull stuck” | Check your internet or try a different mirror |
| “High RAM usage” | Try a smaller model like orca-mini or use quantized versions |
🎓 Exercise
-
Open your terminal
-
Run:
ollama pull mistral -
Then:
ollama run mistral -
Ask the model: “Explain how a car engine works in simple terms.”
✅ Summary
-
Use
ollama pull model-nameto download models -
Run them with
ollama run model-name -
Choose models based on performance and RAM
-
Pulling models is essential before using them in your AI agents
61
