Lesson 3.1: Using Ollama to Run LLMs Locally

Lesson 3.1: Using Ollama to Run LLMs Locally


Lesson Overview

This lesson introduces Ollama, a user-friendly framework and tool designed to run Large Language Models (LLMs) locally on your own machine. You’ll learn how to install Ollama, run pre-trained models, and interact with them without relying on cloud services, which enhances privacy, speed, and control.


Learning Objectives

By the end of this lesson, you will be able to:

  • Understand what Ollama is and its core features.

  • Install Ollama on your local machine (Windows, macOS, or Linux).

  • Download and run popular LLMs using Ollama.

  • Use Ollama’s command-line interface (CLI) to interact with LLMs.

  • Explore basic commands and options for model management.

  • Recognize the benefits and limitations of running LLMs locally with Ollama.


1. Introduction to Ollama

What is Ollama?
Ollama is a lightweight, open-source application designed to simplify the deployment and interaction with large language models locally. It acts as a bridge between you and the LLM, abstracting complex setup and resource management.

Why use Ollama?

  • Privacy: No data leaves your computer.

  • Latency: Responses are immediate with no network delay.

  • Cost: No cloud compute costs or API fees.

  • Control: Full control over model versions and updates.


2. System Requirements

Before installation, ensure your machine meets these minimal requirements:

Component Requirement
OS Windows 10/11, macOS 11+, Linux (Ubuntu/Debian recommended)
CPU Multi-core processor (Intel i5 or better recommended)
RAM At least 16 GB RAM (32 GB or more for larger models)
Disk At least 20 GB free space for models and Ollama installation
GPU Optional but recommended for faster inference (NVIDIA GPU with CUDA support)

3. Installing Ollama

Step 1: Download the installer

  • Visit the official Ollama website: https://ollama.com

  • Download the installer for your operating system.

Step 2: Run the installer

  • Follow the guided prompts to install Ollama.

  • Verify installation by opening a terminal or command prompt and typing:

ollama --version

You should see the installed version printed.


4. Downloading and Running Models with Ollama

Ollama provides easy access to several popular open-source LLMs.

Step 1: List available models

Run:

ollama list

This command shows all models installed locally.

Step 2: Download a model

To download a model, use:

ollama pull <model-name>

Example:

ollama pull llama2

This fetches the LLaMA 2 model locally.

Step 3: Run a model interactively

Start a chat session with the model:

ollama chat llama2

You can now type prompts directly and get responses.


5. Using Ollama CLI to Interact with Models

Some useful Ollama CLI commands:

Command Description
ollama chat <model> Start interactive chat with a model
ollama run <model> Run a single prompt and get output
ollama list List downloaded models
ollama pull <model> Download a model
ollama rm <model> Remove a model from local storage

Example single prompt run:

ollama run llama2 "Explain the theory of relativity in simple terms."

6. Integrating Ollama with Your Applications

  • Ollama supports API access via local HTTP endpoints.

  • You can send prompts programmatically to the local model.

  • Useful for embedding LLMs into local tools, chatbots, or research projects.


7. Best Practices and Tips

  • Start with smaller models for experimentation.

  • Monitor CPU and RAM usage; large models require significant resources.

  • Use GPU acceleration if available for faster inference.

  • Regularly update Ollama and models for latest features and improvements.

  • Keep models offline to ensure privacy and security.


8. Limitations to Keep in Mind

  • Local hardware limits model size and performance.

  • Some models require GPUs for practical speed.

  • Initial downloads can be large and take time.

  • Community support is growing but not as extensive as cloud APIs.


9. Summary

  • Ollama is an accessible tool for running LLMs locally.

  • Installation is straightforward across major OS.

  • You can download, run, and chat with multiple models easily.

  • Ollama provides command-line tools for efficient model management.

  • Running models locally provides privacy, control, and cost benefits.


10. Hands-on Exercise

Task:

  1. Install Ollama on your computer.

  2. Pull the llama2 or any available model.

  3. Start an interactive chat session with the model.

  4. Ask the model three different questions about AI or your topic of interest.

  5. Experiment with ollama run to run single prompts non-interactively.


11. Additional Resources


 

80