Lesson 3.2: Using LM Studio or Text Generation WebUI

 


Lesson 3.2: Using LM Studio or Text Generation WebUI

🎯 Lesson Objectives

By the end of this lesson, learners will be able to:

  • Understand the features and differences between LM Studio and Text Generation WebUI.

  • Install and run either LM Studio or Text Generation WebUI on a local machine.

  • Load a local LLM (like LLaMA, Mistral, or OpenHermes) using either interface.

  • Customize settings such as context size, GPU/CPU usage, and model parameters.

  • Interact with the LLM via a user-friendly interface or API.


🧠 Introduction

Running a large language model (LLM) locally doesn’t require writing your own codebase from scratch. Tools like LM Studio and Text Generation WebUI make it easy to deploy and use powerful open-source models like LLaMA, Mistral, or Gemma on your personal machine. This lesson introduces both platforms and walks you through setup and usage.


🛠️ Part 1: LM Studio

🌟 What is LM Studio?

LM Studio is a polished, user-friendly desktop application that lets you run quantized LLMs (GGUF format) directly on your computer with minimal setup.

✅ Key Features

  • GUI-based interface

  • Auto-download of GGUF models from Hugging Face

  • Chat interface with memory and history

  • Supports GPU acceleration (Apple Silicon, NVIDIA CUDA)

  • No terminal or coding required

📥 Installation

  1. Visit https://lmstudio.ai

  2. Download the installer for your OS (Windows/macOS/Linux).

  3. Run the installer and launch the app.

📦 Loading a Model

  1. Click “Explore Models” to browse Hugging Face GGUF models.

  2. Select a model like TheBloke/Mistral-7B-Instruct-GGUF or OpenHermes-2.5-Mistral-GGUF.

  3. Click “Download” – LM Studio will handle the setup.

  4. Once downloaded, click “Chat” to interact.

⚙️ Settings to Explore

  • Context length (e.g., 4096 tokens)

  • Quantization type (Q4_K_M vs Q6_K)

  • Memory allocation and batch size

  • Model behavior (e.g., temperature, top-k/top-p)

🗣️ Use Case

Perfect for educators, researchers, and developers who want a lightweight, offline chatbot or to experiment with models.


🛠️ Part 2: Text Generation WebUI (Oobabooga)

🌟 What is Text Generation WebUI?

A more advanced, modular interface for running and fine-tuning LLMs. Offers a web interface and plugin system.

✅ Key Features

  • Works with a wide range of backends (GPTQ, GGUF, LLaMA.cpp, etc.)

  • Supports chat, instruct, and roleplay modes

  • LoRA adapter support for fine-tuning

  • API access and multiple user sessions

  • Plugins for voice, image, and code

💻 Installation (Windows/Linux)

Assumes you have Python 3.10+ and Git installed

git clone https://github.com/oobabooga/text-generation-webui.git
cd text-generation-webui
python3 -m venv venv
source venv/bin/activate  # or venvScriptsactivate on Windows
pip install -r requirements.txt

📦 Running the WebUI

python server.py --model TheBloke/Mistral-7B-Instruct-GGUF

You can specify --model-dir and --load-in-8bit to optimize memory usage.

Visit http://localhost:7860 in your browser.

🧩 Add-ons and Extensions

  • AutoGPTQ: For running quantized models with GPU acceleration

  • Character cards and memory: Ideal for roleplay or simulation

  • API Mode: Expose your local LLM as a REST API

⚙️ Settings to Experiment With

  • Sampler settings: temperature, top-k, top-p

  • Prompt formatting: ChatML, Alpaca, Vicuna, etc.

  • GPU/CPU backend selection (CUDA, Metal, ROCm)


🔍 Comparison: LM Studio vs Text Generation WebUI

Feature LM Studio Text Generation WebUI
User Interface GUI desktop app Web browser-based GUI
Installation Complexity Very easy Moderate (requires terminal)
Model Format GGUF only GGUF, GPTQ, safetensors
Backend Flexibility Limited High
Plugin Support No Yes
Ideal For Beginners Power users, tinkerers

🔄 Practice Activity

Task: Download and run Mistral-7B-Instruct in both LM Studio and Text Generation WebUI.
Compare response speed, UI experience, and memory usage. Write a short reflection.


❓ Comprehension Check

  1. What model formats does LM Studio support?

  2. Which tool supports LoRA adapters and multi-user extensions?

  3. How would you optimize a model for low-memory devices in either tool?


📘 Further Learning & Resources


 

111