Lesson 3.3: Hosting on a Cloud Server (e.g., Google Colab, RunPod, Vast.ai)

 


Lesson 3.3: Hosting on a Cloud Server (e.g., Google Colab, RunPod, Vast.ai)


🎯 Lesson Objectives

By the end of this lesson, learners will be able to:

  • Understand the advantages and trade-offs of running LLMs in the cloud.

  • Deploy a local LLM (e.g., Mistral, LLaMA, or GPTQ variants) on platforms like Google Colab, RunPod, and Vast.ai.

  • Choose the right platform for their budget, technical needs, and usage goals.

  • Access their LLM through browser-based interfaces or APIs.

  • Optimize costs, memory, and performance when using cloud infrastructure.


🧠 Introduction

Running LLMs on local machines is great—but it’s not always feasible due to hardware limitations. Cloud platforms like Google Colab, RunPod, and Vast.ai offer GPU-powered environments where you can host models for development, testing, or even production. This lesson walks you through how to get started on each, with examples and key tips.


☁️ Part 1: Google Colab – Quick & Free Cloud GPU

🔍 What is Google Colab?

Google Colab is a free, Jupyter notebook–style environment with access to GPU (and sometimes TPU) resources. Ideal for experimentation, light usage, and prototyping.

✅ Pros

  • Free access to NVIDIA GPUs (T4, A100, or V100 depending on tier)

  • Easy to share notebooks

  • Built-in integration with Google Drive

  • No server maintenance

❌ Cons

  • Time limits (90–120 mins per session)

  • Limited RAM (12–25 GB)

  • No persistent storage unless using Google Drive

⚙️ Steps to Host LLM on Colab

  1. Open Google Colab

  2. Create a new notebook.

  3. Install required dependencies:

!pip install auto-gptq
!pip install transformers accelerate
  1. Load a quantized model (example: Mistral GPTQ from Hugging Face):

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "TheBloke/Mistral-7B-Instruct-v0.1-GPTQ"

tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

prompt = "Explain black holes in simple terms."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
  1. Optional: Connect to Gradio for a simple chatbot UI.


🚀 Part 2: RunPod – Affordable, Easy GPU Hosting

🔍 What is RunPod?

RunPod offers virtualized GPU servers where you can spin up pre-configured environments (called Pods) with persistent storage.

✅ Pros

  • Pay-as-you-go pricing

  • Persistent environment with your code and data

  • Access to fast GPUs like A100, H100

  • Templates for LLMs and inference servers

❌ Cons

  • Requires account and billing setup

  • More technical setup than Colab

⚙️ Steps to Host LLM on RunPod

  1. Visit https://runpod.io and sign up.

  2. Go to “GPU Cloud” → “Deploy a Pod.”

  3. Choose a template like Text Generation WebUI or Docker + Ubuntu.

  4. Select GPU (A100 recommended for larger models).

  5. Once deployed, use the terminal or Jupyter Lab environment to:

    • Clone the Text Generation WebUI

    • Download GGUF or GPTQ models from Hugging Face

    • Run the model with web-based access

  6. Use public IP/port or auto-generated Gradio/streamlit link to interact with the model.


🧠 Part 3: Vast.ai – Flexible, Low-Cost GPU Marketplace

🔍 What is Vast.ai?

Vast.ai is a decentralized marketplace where users rent GPU compute from global providers. Ideal for very cost-sensitive setups.

✅ Pros

  • Cheapest GPU rentals per hour (especially RTX 3090, A6000)

  • Lots of configuration flexibility

  • Choose providers based on location, uptime, and speed

❌ Cons

  • Not beginner-friendly

  • Manual environment setup

  • Less support and documentation

⚙️ Steps to Host LLM on Vast.ai

  1. Visit https://vast.ai and sign up.

  2. Click “Create” > “Container” > Choose oobabooga/text-generation-webui as the image.

  3. Choose machine specs:

    • GPU (3090/4090 or A100 for larger models)

    • RAM and disk (32+ GB RAM and 100 GB SSD ideal)

  4. Once the container starts:

    • Connect via SSH or browser terminal

    • Set up your model:

      git clone https://github.com/oobabooga/text-generation-webui
      cd text-generation-webui
      python3 server.py --model-path [model-path]
      
  5. Access via public IP and port assigned by Vast.


📊 Platform Comparison Table

Feature Google Colab RunPod Vast.ai
Cost Free / Pro $0.10–$2.00/hr $0.05–$1.50/hr
GPU Specs T4, V100, A100 A100, H100, 3090 3090, A6000, A100
Persistent Storage No (unless mounted) Yes Yes
Time Limit Yes No No
Ideal For Beginners, demos Production/test use Advanced, cost-saving
Setup Difficulty Very Easy Moderate Advanced

🧪 Practice Activity

Task: Choose one platform (Google Colab, RunPod, or Vast.ai) and host a quantized model such as TheBloke/Mistral-7B-GGUF.
Goal: Get the model running and chat with it using a basic prompt.
Write down:

  • Time taken to set up

  • GPU type used

  • Cost (if any)

  • Model response time and quality


❓ Comprehension Check

  1. What’s the major limitation of using Google Colab for LLM hosting?

  2. Which platform provides the cheapest hourly GPU cost?

  3. How do you ensure your LLM server remains persistent on RunPod?


📘 Resources


 

80