LLM — Large Language Models

Status: Basic Tags: LLM AI MachineLearning NLP Created: 2026-08-09 Related: Deep Learning, Transformer Architecture, Fine-Tuning

Overview

A Large Language Model (LLM) is an AI model trained on massive text corpora to understand and generate human language. LLMs use neural network architectures, primarily Transformers, to process sequences of text and predict subsequent tokens.

Key Characteristics

  • Scale: Billions to trillions of parameters
  • Training Data: Trillions of tokens from internet text, books, code
  • Capabilities: Text generation, translation, summarization, coding, reasoning
  • Architecture: Transformer-based with self-attention mechanisms

Architecture

Transformer Model

The Transformer architecture (Vaswani et al., 2017) is the foundation of modern LLMs.

Core Components

  1. Self-Attention Mechanism — Models relationships between all tokens in a sequence
  2. Multi-Head Attention — Parallel attention layers capturing different aspects
  3. Position-Wise Feed-Forward — Neural network applied to each position
  4. Layer Normalization & Residual Connections — Stabilize training
  5. Position Encoding — Injects sequence order information

Attention Formula

Attention(Q, K, V) = softmax(QK^T / √d_k)V

Where:

  • Q = Query matrix
  • K = Key matrix
  • V = Value matrix
  • d_k = dimension of key vectors

Variants

  • Decoder-only: GPT series, Llama series — autoregressive generation
  • Encoder-decoder: T5, BERT — understanding and generation
  • Encoder-only: BERT, RoBERTa — text classification, NLU tasks

Parameter Efficiency

ModelParametersArchitectureNotes
GPT-21.5BDecoder-onlyFirst GPT released
GPT-3175BDecoder-onlyBreakthrough scale
LLaMA7-70BDecoder-onlyOpen weight
Mistral7-8x7BDecoder-onlyMoE architecture
Mixtral46.7BMoESparse Mixture of Experts
Claude1T+Decoder-onlyAnthropic’s proprietary
Gemini1.8TDecoder-onlyGoogle’s largest
Qwen110BDecoder-onlyAlibaba’s open model

Training Process

Phases of LLM Training

1. Pre-training (Foundation Model)

  • Goal: Learn general language understanding
  • Data: Internet text, books, Wikipedia, code repositories
  • Objective: Next token prediction (causal language modeling)
  • Compute: Thousands of GPU/TPU hours
  • Duration: Weeks to months

2. Supervised Fine-Tuning (SFT)

  • Goal: Align model with human instructions
  • Data: Human-written prompt-response pairs
  • Method: Continue pre-training on instruction data
  • Result: Model follows instructions reliably

3. Human Feedback

  • RLHF (Reinforcement Learning from Human Feedback):
    • Train reward model on human preferences
    • Optimize LLM with PPO (Proximal Policy Optimization)
  • DPO (Direct Preference Optimization):
    • Direct optimization on preference pairs
    • More stable than RLHF
  • ORPO (Odds Ratio Preference Optimization):
    • Combines SFT and preference optimization

Training Loss

L = -Σ log P(w_t | w_<t; θ)

Key Training Techniques

  • Learning Rate Scheduling: Cosine decay with warmup
  • Batch Size: Global batch size of 2-8M tokens
  • Optimizer: AdamW with β₁=0.9, β₂=0.95
  • Normalization: RMSNorm (more efficient than LayerNorm)
  • Activation: SwiGLU (better performance)

Inference

Inference Methods

Decoding Strategies

StrategyDescriptionQualitySpeed
GreedyAlways pick highest probability tokenLowFastest
Beam SearchKeep top-K sequencesMediumMedium
SamplingPick from probability distributionHighFast
Top-kSample from top-k most likely tokensHighFast
Top-p (nucleus)Sample from cumulative probability pHighFast
TemperatureControl randomness (higher = more random)VariableFast

Optimization Techniques

  • KV Cache: Cache key-value pairs for efficient generation
  • Quantization: FP16, INT8, INT4, GPTQ, AWQ
  • Speculative Decoding: Draft model + verify model
  • FlashAttention: Optimized attention implementation
  • PagedAttention: Efficient memory management (vLLM)

Hardware Requirements

Model SizeFP16 VRAMINT4 VRAMRecommended GPU
7B~14GB~5GBRTX 3090/4090
13B~26GB~10GBRTX 4090
30B~60GB~20GBA100 80GB
70B~140GB~40GBA100 80GB × 2
405B~810GB~200GBA100 80GB × 8+

Applications

Common Use Cases

  • Chatbots: Conversational AI assistants
  • Code Generation: GitHub Copilot, Cursor
  • Document Analysis: Summarization, extraction
  • Translation: Machine translation between languages
  • Content Creation: Marketing, articles, emails
  • Research: Literature review, hypothesis generation
  • Education: Tutoring, personalized learning

Prompt Engineering Techniques

  • Zero-shot: Direct prompt without examples
  • One-shot: Single example in prompt
  • Few-shot: Multiple examples in prompt
  • Chain-of-Thought: “Let’s think step by step”
  • ReAct: Reasoning + acting pattern
  • Tree of Thoughts: Explore multiple reasoning paths
  • Self-Consistency: Sample multiple paths, vote on answer

Limitations & Challenges

Known Issues

  • Hallucination: Confidently generates incorrect information
  • Context Window Limits: Finite token limit for input
  • Bias: Reflects biases in training data
  • Copyright: Training data raises legal concerns
  • Compute Cost: Massive resources required for training
  • Safety: Potential for misuse (deepfakes, misinformation)
  • Reasoning: Struggles with complex logical deduction
  • Factual Accuracy: Outdated knowledge beyond training data

Mitigation Strategies

  • Retrieval-Augmented Generation (RAG): Augment with external knowledge
  • Constitutional AI: Self-correction via principles
  • Red-teaming: Adversarial testing before deployment
  • Content Filtering: Safety classifiers
  • Fact-checking: Cross-reference with reliable sources

Key Papers & Resources

Foundational Papers

  1. “Attention Is All You Need” — Vaswani et al. (2017)
  2. “BERT: Pre-training of Deep Bidirectional Transformers” — Devlin et al. (2019)
  3. “GPT-2: Language Models are Unsupervised Multitask Learners” — Radford et al. (2019)
  4. “GPT-3: Language Models are Few-Shot Learners” — Brown et al. (2020)
  5. “LLaMA: Open and Efficient Foundation Language Models” — Touvron et al. (2023)
  6. “LoRA: Low-Rank Adaptation of Large Language Models” — Hu et al. (2021)
  7. “FlashAttention: Fast and Memory-Efficient Exact Attention” — Dao et al. (2022)

Learning Resources