Flashcards
Spaced repetition cards. Requires Obsidian Spaced Repetition plugin. Format: Front\n\nBack
Models
What is a Large Language Model (LLM)?
LLM = Neural network trained on massive text corpora to predict next token. Key architectures: Transformer (decoder-only, e.g., GPT), Encoder-Decoder (e.g., T5). Measured by parameter count, context window, and benchmark performance.
What is the Transformer architecture?
Core architecture for modern LLMs (Vaswani et al., 2017). Key components: Self-attention mechanism, multi-head attention, positional encoding, feed-forward layers, residual connections. Replaces RNNs/CNNs with parallelizable attention over all positions.
What is prompt engineering?
Techniques for crafting inputs that elicit desired LLM behavior: few-shot learning, chain-of-thought, role-playing, structured outputs, temperature tuning, system prompts, delimiters, and negative prompting.
What is RAG (Retrieval-Augmented Generation)?
Architecture combining retrieval with generation: retrieve relevant documents from vector DB → concatenate with query → pass to LLM for grounded generation. Solves LLM limitations: outdated knowledge, hallucination, private data access.
What is quantization?
Reducing model precision to shrink size/increase speed: FP16 (half), INT8 (8-bit), INT4 (4-bit), GGUF formats. Trade-off: ~10-20% quality loss for 2-8x speedup. Tools: llama.cpp, bitsandbytes, AWQ.
What are MoE (Mixture of Experts) models?
Architecture using multiple “expert” sub-networks, with a router selecting which experts to activate per token. Benefits: larger total parameters, lower compute per forward pass. Examples: Mixtral 8x7B, Grok-1.
What is RLHF (Reinforcement Learning from Human Feedback)?
Training technique: train reward model from human rankings → use PPO to fine-tune LLM to maximize reward. Purpose: align model output with human preferences (helpful, honest, harmless). Alternatives: DPO, ORPO, RLVR.
What is context window?
Number of tokens an LLM can process in a single prompt+response. Modern models: 128K-1M+ tokens. Larger windows enable full-document reasoning but increase compute O(n²) due to attention. Techniques: sliding window, ring attention, positional interpolation.
What is embedding/vector representation?
Dense numerical vector that encodes semantic meaning of text. Used for: similarity search, RAG, clustering, transfer learning. Measured by cosine similarity or dot product. Models: text-embedding-3-small (OpenAI), BGE (Tencent), Nomic Embed.
What is fine-tuning vs. prompting?
Fine-tuning: modify model weights on task-specific data (expensive, persistent, requires GPU). Prompting: craft input to elicit behavior (cheap, no weights change, per-request). Best practice: use prompting for most tasks, fine-tune only when prompting fails.
What is an attention mechanism?
Core Transformer component that computes weighted relationships between all tokens in a sequence: Attention(Q,K,V) = softmax(QKᵀ/√dₖ)V. Enables the model to focus on relevant parts of input regardless of distance.
What is a tokenizer?
Component that converts text → tokens (integers) before LLM processing. Types: Byte-Pair Encoding (BPE), SentencePiece, Unigram. Vocabulary size affects: token count, context window efficiency, language support.
What is the hallucination problem?
LLMs generating plausible but incorrect/fabricated information. Causes: training on noise, overgeneralization, lack of grounding. Mitigations: RAG, citation requirements, confidence scoring, fact-checking layers.
What is parameter-efficient fine-tuning (PEFT)?
Techniques to fine-tune large models with minimal parameter changes: LoRA (low-rank adaptation), QLoRA (quantized LoRA), adapters, prefix tuning. Enables fine-tuning models on consumer GPUs.
What is a synthetic dataset?
Artificially generated training data: use stronger model to create data for weaker model (e.g., DistilBERT, OLMo). Advantages: unlimited scale, clean labels, targeted coverage. Risks: model collapse, error propagation, bias amplification.
What is the scaling law in LLMs?
Emergent finding: model performance scales predictably with compute, data, and parameters (Chinchilla scaling). Optimal training uses more data and parameters for same compute budget. Still holds at extreme scale (100T+ parameters).
What is an AI agent?
System that uses an LLM as a reasoning engine to plan and execute multi-step tasks: tool use (API calls, code execution), memory (short-term context, long-term vector DB), planning (reAct, ToT, GoT), multi-agent coordination.
What is the difference between a model and an API?
Model: raw weights you run yourself (full control, privacy, self-hosted, requires GPU). API: hosted service (easier, pay-per-use, data leaves your machine, rate limits, less control). Examples: OpenAI API vs. running Llama locally.
What is multimodal AI?
Models that process multiple input types: text, images, audio, video, 3D. Architecture: shared encoder, cross-attention between modalities. Examples: GPT-4V (vision+text), Whisper (audio→text), LLaVA (vision-language).