✅ Lesson 3.4: Loading and Testing Local Models

 


✅ Lesson 3.4: Loading and Testing Local Models


🎯 Lesson Objectives

By the end of this lesson, you will:

  • Understand how to prepare and load different types of LLM model files.

  • Know how to run models using LM Studio, Text Generation WebUI, or CLI-based tools.

  • Interact with LLMs using structured prompts and evaluate their output.

  • Troubleshoot common loading errors and optimize performance.

  • Conduct performance testing (response time, memory usage, coherence).


🧠 1. Why Loading & Testing Matters

Even with the best interface or environment, a model is only useful once it’s properly loaded, tested, and verified. Local loading allows you to:

  • Keep data private (no internet required).

  • Customize usage (fine-tuned prompts, system instructions).

  • Benchmark performance on your own hardware.

  • Understand model limits (context length, speed, RAM/GPU requirements).


🧰 2. Choosing a Model Format

Before you load a model, select a format that matches your system specs and preferred tools:

Format Extension For Use With Ideal When…
GGUF .gguf llama.cpp, LM Studio, WebUI You want quantized, fast CPU/GPU
GPTQ .safetensors Transformers + AutoGPTQ You have a GPU and want speed
FP16 .bin or .pt Transformers (full precision) You have 24GB+ VRAM and accuracy

➡️ Recommended repos: TheBloke, NousResearch, mistralai


🖥️ 3. Loading a Model: Method-by-Method


🔹 Option A: Using LM Studio (for GGUF models)

  1. Download LM Studio: https://lmstudio.ai

  2. Open the App > Click “Explore models”.

  3. Search for models like Mistral, OpenHermes, or Gemma.

  4. Download the GGUF version (e.g., Q4_K_M.gguf).

  5. Click Chat, type a prompt like:

    Explain what a neural network is in simple terms.
    
  6. Adjust temperature, top-k, and context length in the settings tab.

💡 LM Studio auto-detects Apple Silicon or CUDA if available.


🔹 Option B: Using Text Generation WebUI (GUI + flexible backends)

  1. Install it (if not already):

    git clone https://github.com/oobabooga/text-generation-webui
    cd text-generation-webui
    python3 -m venv venv
    source venv/bin/activate
    pip install -r requirements.txt
    
  2. Download your model (put in /models folder). Example:

    • Mistral-7B-Instruct-v0.1.Q4_K_M.gguf

    • TheBloke/OpenHermes-2.5-GPTQ

  3. Run the WebUI:

    python3 server.py --model mistral-7b-instruct-v0.1.Q4_K_M.gguf
    
  4. Open browser: http://localhost:7860


🔹 Option C: Using CLI (llama.cpp)

  1. Build the tool:

    git clone https://github.com/ggerganov/llama.cpp
    cd llama.cpp
    make
    
  2. Run a model:

    ./main -m ./models/mistral-7b-instruct-v0.1.Q4_K_M.gguf -p "What is quantum computing?"
    
  3. Optional flags:

    • -t 8 for thread count

    • -c 2048 for context tokens

    • --color for colorful output


🧪 4. Testing Prompts

Use diverse prompt types to test the model’s capabilities:

📚 Knowledge & Factual Retrieval

What are the three laws of thermodynamics?

✍️ Creative Generation

Write a short story about a robot that becomes a poet.

💼 Instruction-Following

Summarize this: "Large Language Models are trained by..." (add a real paragraph)

🌍 Multilingual

Translate into French: “How are you today?”

🧠 Role-Based Prompting

You are a math teacher. Explain the Pythagorean Theorem to a 12-year-old.

📈 5. Evaluating Performance

After loading and running a few prompts, evaluate:

Metric Method
Latency Time from prompt to response (use timer or stopwatch)
Output Quality Coherence, fluency, relevance
Memory Usage Use htop, nvidia-smi, or Activity Monitor
Token Limit Try longer context prompts and see when cutoff occurs
Prompt Sensitivity Slight prompt changes – do answers stay consistent?

🛠️ 6. Troubleshooting Guide

Issue Solution
Model doesn’t load Check RAM/GPU availability; reduce model size (Q4_K_M or Q5_1)
Output is gibberish Wrong tokenizer or missing system instruction prompt formatting
Slow generation Lower context size or use quantized models
Error: CUDA out of memory Try --load-in-8bit or switch to smaller model
Model freezes at load Use --no-stream or check compatibility between model and UI

🔄 7. Practice Activity

🧪 Assignment:

  1. Download two different models:

    • One in GGUF (e.g., Mistral-7B-Q4_K_M)

    • One in GPTQ (e.g., OpenHermes-2.5-GPTQ)

  2. Load each using two different tools (LM Studio and WebUI or CLI).

  3. Prompt: “Write a motivational message for someone starting a new job.”

  4. Record:

    • Load time

    • Output fluency

    • Memory usage

    • Subjective quality (1–5 stars)


❓ 8. Comprehension Check

  1. What’s the main difference between a GGUF and GPTQ model?

  2. What’s one reason a model might give incoherent output even if it loads correctly?

  3. How do you reduce memory usage when running a model locally?


📘 9. Further Learning & Resources

 

91