Lesson 3.3: Hosting on a Cloud Server (e.g., Google Colab, RunPod, Vast.ai)
🎯 Lesson Objectives
By the end of this lesson, learners will be able to:
-
Understand the advantages and trade-offs of running LLMs in the cloud.
-
Deploy a local LLM (e.g., Mistral, LLaMA, or GPTQ variants) on platforms like Google Colab, RunPod, and Vast.ai.
-
Choose the right platform for their budget, technical needs, and usage goals.
-
Access their LLM through browser-based interfaces or APIs.
-
Optimize costs, memory, and performance when using cloud infrastructure.
🧠 Introduction
Running LLMs on local machines is great—but it’s not always feasible due to hardware limitations. Cloud platforms like Google Colab, RunPod, and Vast.ai offer GPU-powered environments where you can host models for development, testing, or even production. This lesson walks you through how to get started on each, with examples and key tips.
☁️ Part 1: Google Colab – Quick & Free Cloud GPU
🔍 What is Google Colab?
Google Colab is a free, Jupyter notebook–style environment with access to GPU (and sometimes TPU) resources. Ideal for experimentation, light usage, and prototyping.
✅ Pros
-
Free access to NVIDIA GPUs (T4, A100, or V100 depending on tier)
-
Easy to share notebooks
-
Built-in integration with Google Drive
-
No server maintenance
❌ Cons
-
Time limits (90–120 mins per session)
-
Limited RAM (12–25 GB)
-
No persistent storage unless using Google Drive
⚙️ Steps to Host LLM on Colab
-
Open Google Colab
-
Create a new notebook.
-
Install required dependencies:
!pip install auto-gptq
!pip install transformers accelerate
-
Load a quantized model (example: Mistral GPTQ from Hugging Face):
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "TheBloke/Mistral-7B-Instruct-v0.1-GPTQ"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
prompt = "Explain black holes in simple terms."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
-
Optional: Connect to Gradio for a simple chatbot UI.
🚀 Part 2: RunPod – Affordable, Easy GPU Hosting
🔍 What is RunPod?
RunPod offers virtualized GPU servers where you can spin up pre-configured environments (called Pods) with persistent storage.
✅ Pros
-
Pay-as-you-go pricing
-
Persistent environment with your code and data
-
Access to fast GPUs like A100, H100
-
Templates for LLMs and inference servers
❌ Cons
-
Requires account and billing setup
-
More technical setup than Colab
⚙️ Steps to Host LLM on RunPod
-
Visit https://runpod.io and sign up.
-
Go to “GPU Cloud” → “Deploy a Pod.”
-
Choose a template like
Text Generation WebUIorDocker + Ubuntu. -
Select GPU (A100 recommended for larger models).
-
Once deployed, use the terminal or Jupyter Lab environment to:
-
Clone the Text Generation WebUI
-
Download GGUF or GPTQ models from Hugging Face
-
Run the model with web-based access
-
-
Use public IP/port or auto-generated Gradio/streamlit link to interact with the model.
🧠 Part 3: Vast.ai – Flexible, Low-Cost GPU Marketplace
🔍 What is Vast.ai?
Vast.ai is a decentralized marketplace where users rent GPU compute from global providers. Ideal for very cost-sensitive setups.
✅ Pros
-
Cheapest GPU rentals per hour (especially RTX 3090, A6000)
-
Lots of configuration flexibility
-
Choose providers based on location, uptime, and speed
❌ Cons
-
Not beginner-friendly
-
Manual environment setup
-
Less support and documentation
⚙️ Steps to Host LLM on Vast.ai
-
Visit https://vast.ai and sign up.
-
Click “Create” > “Container” > Choose
oobabooga/text-generation-webuias the image. -
Choose machine specs:
-
GPU (3090/4090 or A100 for larger models)
-
RAM and disk (32+ GB RAM and 100 GB SSD ideal)
-
-
Once the container starts:
-
Connect via SSH or browser terminal
-
Set up your model:
git clone https://github.com/oobabooga/text-generation-webui cd text-generation-webui python3 server.py --model-path [model-path]
-
-
Access via public IP and port assigned by Vast.
📊 Platform Comparison Table
| Feature | Google Colab | RunPod | Vast.ai |
|---|---|---|---|
| Cost | Free / Pro | $0.10–$2.00/hr | $0.05–$1.50/hr |
| GPU Specs | T4, V100, A100 | A100, H100, 3090 | 3090, A6000, A100 |
| Persistent Storage | No (unless mounted) | Yes | Yes |
| Time Limit | Yes | No | No |
| Ideal For | Beginners, demos | Production/test use | Advanced, cost-saving |
| Setup Difficulty | Very Easy | Moderate | Advanced |
🧪 Practice Activity
Task: Choose one platform (Google Colab, RunPod, or Vast.ai) and host a quantized model such as
TheBloke/Mistral-7B-GGUF.
Goal: Get the model running and chat with it using a basic prompt.
Write down:
Time taken to set up
GPU type used
Cost (if any)
Model response time and quality
❓ Comprehension Check
-
What’s the major limitation of using Google Colab for LLM hosting?
-
Which platform provides the cheapest hourly GPU cost?
-
How do you ensure your LLM server remains persistent on RunPod?
📘 Resources
80
