Flashcards

Spaced repetition cards. Requires Obsidian Spaced Repetition plugin. Format: Front\n\nBack

Models

What is a Large Language Model (LLM)?

LLM = Neural network trained on massive text corpora to predict next token. Key architectures: Transformer (decoder-only, e.g., GPT), Encoder-Decoder (e.g., T5). Measured by parameter count, context window, and benchmark performance.


What is the Transformer architecture?

Core architecture for modern LLMs (Vaswani et al., 2017). Key components: Self-attention mechanism, multi-head attention, positional encoding, feed-forward layers, residual connections. Replaces RNNs/CNNs with parallelizable attention over all positions.


What is prompt engineering?

Techniques for crafting inputs that elicit desired LLM behavior: few-shot learning, chain-of-thought, role-playing, structured outputs, temperature tuning, system prompts, delimiters, and negative prompting.


What is RAG (Retrieval-Augmented Generation)?

Architecture combining retrieval with generation: retrieve relevant documents from vector DB → concatenate with query → pass to LLM for grounded generation. Solves LLM limitations: outdated knowledge, hallucination, private data access.


What is quantization?

Reducing model precision to shrink size/increase speed: FP16 (half), INT8 (8-bit), INT4 (4-bit), GGUF formats. Trade-off: ~10-20% quality loss for 2-8x speedup. Tools: llama.cpp, bitsandbytes, AWQ.


What are MoE (Mixture of Experts) models?

Architecture using multiple “expert” sub-networks, with a router selecting which experts to activate per token. Benefits: larger total parameters, lower compute per forward pass. Examples: Mixtral 8x7B, Grok-1.


What is RLHF (Reinforcement Learning from Human Feedback)?

Training technique: train reward model from human rankings → use PPO to fine-tune LLM to maximize reward. Purpose: align model output with human preferences (helpful, honest, harmless). Alternatives: DPO, ORPO, RLVR.


What is context window?

Number of tokens an LLM can process in a single prompt+response. Modern models: 128K-1M+ tokens. Larger windows enable full-document reasoning but increase compute O(n²) due to attention. Techniques: sliding window, ring attention, positional interpolation.


What is embedding/vector representation?

Dense numerical vector that encodes semantic meaning of text. Used for: similarity search, RAG, clustering, transfer learning. Measured by cosine similarity or dot product. Models: text-embedding-3-small (OpenAI), BGE (Tencent), Nomic Embed.


What is fine-tuning vs. prompting?

Fine-tuning: modify model weights on task-specific data (expensive, persistent, requires GPU). Prompting: craft input to elicit behavior (cheap, no weights change, per-request). Best practice: use prompting for most tasks, fine-tune only when prompting fails.


What is an attention mechanism?

Core Transformer component that computes weighted relationships between all tokens in a sequence: Attention(Q,K,V) = softmax(QKᵀ/√dₖ)V. Enables the model to focus on relevant parts of input regardless of distance.


What is a tokenizer?

Component that converts text → tokens (integers) before LLM processing. Types: Byte-Pair Encoding (BPE), SentencePiece, Unigram. Vocabulary size affects: token count, context window efficiency, language support.


What is the hallucination problem?

LLMs generating plausible but incorrect/fabricated information. Causes: training on noise, overgeneralization, lack of grounding. Mitigations: RAG, citation requirements, confidence scoring, fact-checking layers.


What is parameter-efficient fine-tuning (PEFT)?

Techniques to fine-tune large models with minimal parameter changes: LoRA (low-rank adaptation), QLoRA (quantized LoRA), adapters, prefix tuning. Enables fine-tuning models on consumer GPUs.


What is a synthetic dataset?

Artificially generated training data: use stronger model to create data for weaker model (e.g., DistilBERT, OLMo). Advantages: unlimited scale, clean labels, targeted coverage. Risks: model collapse, error propagation, bias amplification.


What is the scaling law in LLMs?

Emergent finding: model performance scales predictably with compute, data, and parameters (Chinchilla scaling). Optimal training uses more data and parameters for same compute budget. Still holds at extreme scale (100T+ parameters).


What is an AI agent?

System that uses an LLM as a reasoning engine to plan and execute multi-step tasks: tool use (API calls, code execution), memory (short-term context, long-term vector DB), planning (reAct, ToT, GoT), multi-agent coordination.


What is the difference between a model and an API?

Model: raw weights you run yourself (full control, privacy, self-hosted, requires GPU). API: hosted service (easier, pay-per-use, data leaves your machine, rate limits, less control). Examples: OpenAI API vs. running Llama locally.


What is multimodal AI?

Models that process multiple input types: text, images, audio, video, 3D. Architecture: shared encoder, cross-attention between modalities. Examples: GPT-4V (vision+text), Whisper (audio→text), LLaVA (vision-language).