✅ Lesson 3.4: Loading and Testing Local Models
🎯 Lesson Objectives
By the end of this lesson, you will:
-
Understand how to prepare and load different types of LLM model files.
-
Know how to run models using LM Studio, Text Generation WebUI, or CLI-based tools.
-
Interact with LLMs using structured prompts and evaluate their output.
-
Troubleshoot common loading errors and optimize performance.
-
Conduct performance testing (response time, memory usage, coherence).
🧠 1. Why Loading & Testing Matters
Even with the best interface or environment, a model is only useful once it’s properly loaded, tested, and verified. Local loading allows you to:
-
Keep data private (no internet required).
-
Customize usage (fine-tuned prompts, system instructions).
-
Benchmark performance on your own hardware.
-
Understand model limits (context length, speed, RAM/GPU requirements).
🧰 2. Choosing a Model Format
Before you load a model, select a format that matches your system specs and preferred tools:
| Format | Extension | For Use With | Ideal When… |
|---|---|---|---|
| GGUF | .gguf |
llama.cpp, LM Studio, WebUI | You want quantized, fast CPU/GPU |
| GPTQ | .safetensors |
Transformers + AutoGPTQ | You have a GPU and want speed |
| FP16 | .bin or .pt |
Transformers (full precision) | You have 24GB+ VRAM and accuracy |
➡️ Recommended repos: TheBloke, NousResearch, mistralai
🖥️ 3. Loading a Model: Method-by-Method
🔹 Option A: Using LM Studio (for GGUF models)
-
Download LM Studio: https://lmstudio.ai
-
Open the App > Click “Explore models”.
-
Search for models like
Mistral,OpenHermes, orGemma. -
Download the GGUF version (e.g.,
Q4_K_M.gguf). -
Click Chat, type a prompt like:
Explain what a neural network is in simple terms. -
Adjust temperature, top-k, and context length in the settings tab.
💡 LM Studio auto-detects Apple Silicon or CUDA if available.
🔹 Option B: Using Text Generation WebUI (GUI + flexible backends)
-
Install it (if not already):
git clone https://github.com/oobabooga/text-generation-webui cd text-generation-webui python3 -m venv venv source venv/bin/activate pip install -r requirements.txt -
Download your model (put in
/modelsfolder). Example:-
Mistral-7B-Instruct-v0.1.Q4_K_M.gguf -
TheBloke/OpenHermes-2.5-GPTQ
-
-
Run the WebUI:
python3 server.py --model mistral-7b-instruct-v0.1.Q4_K_M.gguf -
Open browser:
http://localhost:7860
🔹 Option C: Using CLI (llama.cpp)
-
Build the tool:
git clone https://github.com/ggerganov/llama.cpp cd llama.cpp make -
Run a model:
./main -m ./models/mistral-7b-instruct-v0.1.Q4_K_M.gguf -p "What is quantum computing?" -
Optional flags:
-
-t 8for thread count -
-c 2048for context tokens -
--colorfor colorful output
-
🧪 4. Testing Prompts
Use diverse prompt types to test the model’s capabilities:
📚 Knowledge & Factual Retrieval
What are the three laws of thermodynamics?
✍️ Creative Generation
Write a short story about a robot that becomes a poet.
💼 Instruction-Following
Summarize this: "Large Language Models are trained by..." (add a real paragraph)
🌍 Multilingual
Translate into French: “How are you today?”
🧠 Role-Based Prompting
You are a math teacher. Explain the Pythagorean Theorem to a 12-year-old.
📈 5. Evaluating Performance
After loading and running a few prompts, evaluate:
| Metric | Method |
|---|---|
| Latency | Time from prompt to response (use timer or stopwatch) |
| Output Quality | Coherence, fluency, relevance |
| Memory Usage | Use htop, nvidia-smi, or Activity Monitor |
| Token Limit | Try longer context prompts and see when cutoff occurs |
| Prompt Sensitivity | Slight prompt changes – do answers stay consistent? |
🛠️ 6. Troubleshooting Guide
| Issue | Solution |
|---|---|
| Model doesn’t load | Check RAM/GPU availability; reduce model size (Q4_K_M or Q5_1) |
| Output is gibberish | Wrong tokenizer or missing system instruction prompt formatting |
| Slow generation | Lower context size or use quantized models |
| Error: CUDA out of memory | Try --load-in-8bit or switch to smaller model |
| Model freezes at load | Use --no-stream or check compatibility between model and UI |
🔄 7. Practice Activity
🧪 Assignment:
Download two different models:
One in GGUF (e.g., Mistral-7B-Q4_K_M)
One in GPTQ (e.g., OpenHermes-2.5-GPTQ)
Load each using two different tools (LM Studio and WebUI or CLI).
Prompt: “Write a motivational message for someone starting a new job.”
Record:
Load time
Output fluency
Memory usage
Subjective quality (1–5 stars)
❓ 8. Comprehension Check
-
What’s the main difference between a GGUF and GPTQ model?
-
What’s one reason a model might give incoherent output even if it loads correctly?
-
How do you reduce memory usage when running a model locally?
📘 9. Further Learning & Resources
91
