AI LLM Keyword Map
Cách Đọc
Map này giữ canonical English spelling. Mục tiêu là keyword để tra lại, không phải glossary.
Landmark model = model hoặc hệ đã tạo cú hích rõ trong cộng đồng AI/LLM, sản phẩm, open weights, agent, multimodal, reasoning, video/audio, hoặc infra.
Nền Tảng
- Artificial Intelligence, Machine Learning, Deep Learning, Generative AI, Foundation Model, Large Language Model, Vision-Language Model, Multimodal Large Language Model, Large Action Model, Frontier Model, Open-Weight Model, Open-Source Model, Proprietary Model, Local LLM, On-Device AI, Edge AI, Sovereign AI
- Token, tokenizer, vocabulary, byte pair encoding, SentencePiece, token budget, context window, long context, completion, logit, logprob, decoding, temperature, top-k, top-p, nucleus sampling, beam search, stop sequence, seed, deterministic decoding
- Embedding, vector, latent space, representation, manifold, semantic similarity, cosine similarity, embedding model, cross-encoder, bi-encoder, reranker, sparse retrieval, dense retrieval, hybrid retrieval
- Emergence, scaling laws, compute-optimal training, Chinchilla scaling, data quality, data mixture, synthetic data, data contamination, benchmark saturation, distribution shift, grokking, phase transition, capability overhang
Transformer Stack
- Transformer, decoder-only Transformer, encoder-decoder Transformer, encoder-only Transformer, attention, self-attention, cross-attention, causal attention, masked attention, multi-head attention, multi-query attention, grouped-query attention, multi-head latent attention, sliding-window attention, block-sparse attention, linear attention
- Query, Key, Value, QKV projection, attention score, attention head, attention sink, KV cache, paged KV cache, prefix cache, prompt cache, rotary position embedding, ALiBi, absolute position embedding, relative position embedding
- residual stream, residual connection, layer norm, RMSNorm, pre-norm, post-norm, MLP block, feed-forward network, SwiGLU, GeGLU, activation function, softmax, logits processor
- Mixture of Experts, sparse MoE, dense model, router, expert, top-k routing, expert parallelism, load balancing, auxiliary loss, auxiliary-loss-free load balancing, DeepSeekMoE, Switch Transformer, Mixtral, active parameters, total parameters
- State Space Model, Mamba, RetNet, RWKV, Hyena, Titans, memory layer, recurrent memory, learned meta-tokens, compression token, context compression
Training
- pretraining, continual pretraining, post-training, supervised fine-tuning, instruction tuning, chat tuning, preference tuning, reinforcement learning from human feedback, reinforcement learning from AI feedback, Constitutional AI, Direct Preference Optimization, Kahneman-Tversky Optimization, Proximal Policy Optimization, Group Relative Policy Optimization, reinforcement learning with verifiable rewards
- next-token prediction, multi-token prediction, masked language modeling, denoising objective, contrastive learning, joint-embedding predictive architecture, curriculum learning, cold-start data, rejection sampling, self-play, expert iteration, distillation, logit distillation, reasoning distillation
- LoRA, QLoRA, AdaLoRA, IA3, PEFT, full fine-tuning, adapter, prompt tuning, prefix tuning, quantization-aware training, pruning, sparsity, weight tying, checkpoint averaging
- dataset curation, deduplication, filtering, toxicity filtering, PII filtering, copyright filtering, data provenance, data flywheel, weak supervision, human annotation, preference data, synthetic reasoning trace, tool-use trace, trajectory data
- chain of thought, hidden reasoning, visible reasoning, process supervision, outcome supervision, verifier, reward model, critique model, debate, tree of thoughts, graph of thoughts, inference-time scaling, deliberation, reflection
Inference Và Serving
- prefill, decode, time to first token, tokens per second, throughput, latency, batch size, continuous batching, dynamic batching, speculative decoding, draft model, target model, Medusa, Lookahead decoding, early exit
- FlashAttention, FlashAttention-2, FlashAttention-3, PagedAttention, vLLM, TensorRT-LLM, llama.cpp, GGUF, Ollama, MLX, SGLang, TGI, Triton, CUDA, ROCm, TPU, XLA, H100, H200, B200, GB200, Blackwell, Trainium, Inferentia
- FP32, BF16, FP16, FP8, FP4, INT8, INT4, NF4, GPTQ, AWQ, SmoothQuant, KV-cache quantization, activation quantization, weight-only quantization
- model routing, mixture of models, cascade, fallback model, router model, cost-aware routing, latency-aware routing, context caching, cache read, cache write, Batch API, streaming, server-sent events, structured outputs
Prompt Và Context Engineering
- prompt engineering, context engineering, system prompt, developer message, user message, instruction hierarchy, few-shot prompting, zero-shot prompting, negative prompting, prompt template, metaprompt, prompt chaining
- retrieval context, working memory, episodic memory, semantic memory, procedural memory, project memory, scratchpad memory, conversation summary, memory compaction, context pruning, context packing, context window management
- indirect prompt injection, instruction conflict, context poisoning, tool-result poisoning, prompt leakage, system prompt extraction, DAN, many-shot jailbreak, universal adversarial trigger, refusal bypass, policy sandwich
RAG Và Knowledge Systems
- Retrieval-Augmented Generation, GraphRAG, agentic RAG, corrective RAG, self-RAG, adaptive RAG, long-context RAG, query rewriting, query decomposition, multi-hop retrieval, citation grounding
- chunking, semantic chunking, parent-child chunking, sliding-window chunking, recursive splitter, metadata filter, document loader, vector database, FAISS, Milvus, Weaviate, Pinecone, Chroma, pgvector, LanceDB, Vespa, Elasticsearch, OpenSearch, BM25
- embedding index, ANN search, HNSW, IVF, product quantization, reranking, cross-encoder reranker, ColBERT, SPLADE, HyDE, fusion retrieval, reciprocal rank fusion
- knowledge graph, entity extraction, ontology, schema, provenance, freshness, source authority, answerability, abstention, hallucination check, citation check
Agents, Tools, Protocols
- agent, AI agent, autonomous agent, semi-autonomous agent, managed agent, browser agent, coding agent, computer-use agent, research agent, data agent, voice agent, embodied agent
- harness, harness engineering, orchestration, planner-executor, supervisor, worker agent, multi-agent system, swarm, delegation, handoff, tool loop, observe-think-act, ReAct, plan-and-execute, reflection loop, evaluator-optimizer
- tool calling, function calling, custom tool, hosted tool, code interpreter, file search, web search, computer use, browser use, shell tool, apply patch tool, structured output, JSON schema, strict schema, tool schema, tool result, tool trace
- Model Context Protocol, MCP server, MCP client, MCP host, MCP transport, resource, prompt, tool, sampling, elicitation, roots, authorization, stateless MCP core, MCP Apps
- Agent2Agent Protocol, A2A, AgentCard, agent discovery, agent-to-agent messaging, task handoff, remote agent, interoperability, Agentic AI Foundation
- skill, plugin, connector, app, workflow, automation, trigger, guardrail, policy check, approval gate, audit log, trace, telemetry
Alignment, Safety, Security
- alignment, helpfulness, harmlessness, honesty, corrigibility, robustness, red teaming, adversarial testing, frontier safety, preparedness framework, responsible scaling policy, safety case
- guardrails, input guardrail, output guardrail, content policy, refusal, over-refusal, under-refusal, safe completion, crisis routing, abuse monitoring, policy model, moderation model
- hallucination, confabulation, sycophancy, deception, sandbagging, sleeper agent, scheming, situational awareness, reward hacking, specification gaming, goal misgeneralization, power-seeking, shutdown resistance
- jailbreak, prompt injection, data exfiltration, tool abuse, SSRF via tool, OAuth consent attack, malicious MCP server, poisoned connector, sandbox escape, cyber capability, biosecurity, CBRN, model weight theft, eval leakage
- model card, system card, transparency report, audit, external evaluator, METR, Apollo Research, ARC Evals, Frontier Safety Framework
Evals Và Benchmarks
- MMLU, MMLU-Pro, GPQA, GPQA Diamond, GSM8K, MATH, MATH-500, AIME, FrontierMath, Humanity's Last Exam, BIG-bench, HELM, MMMU, MathVista
- HumanEval, MBPP, LiveCodeBench, SWE-bench, SWE-bench Verified, SWE-rebench, Aider Polyglot, Terminal-Bench, Terminal-Bench-Science, Codeforces rating, LeetCode benchmark
- ARC-AGI, ARC-AGI-2, ARC-AGI-3, Abstraction and Reasoning Corpus, OSWorld, WebArena, BrowserGym, WorkArena, AndroidWorld, τ-bench, TAU2-Bench, AutomationBench
- Chatbot Arena, LMSYS Arena, Elo, win rate, pairwise preference, LLM-as-judge, judge model, rubric, gold label, pass@k, contamination check, holdout set, private eval
Multimodal
- multimodal input, multimodal output, omni model, native multimodal, any-to-any model, early fusion, late fusion, modality adapter, cross-modal attention
- vision encoder, image encoder, patch embedding, OCR, visual grounding, object detection, segmentation, SAM, grounding, visual question answering, document understanding, chart understanding, screenshot understanding
- speech-to-text, text-to-speech, speech-to-speech, full-duplex speech, realtime audio, turn detection, voice activity detection, prosody, speaker diarization, voice cloning, audio token, codec model
- video understanding, video generation, text-to-video, image-to-video, video-to-video, temporal consistency, object permanence, scene dynamics, camera control, lip sync, synchronized audio, storyboard, shot, keyframe
- 3D generation, NeRF, Gaussian splatting, text-to-3D, robotics perception, sensor fusion
Diffusion, Image, Video, Audio
- diffusion model, score-based model, denoising diffusion probabilistic model, DDIM, latent diffusion, rectified flow, flow matching, consistency model, diffusion transformer, DiT, U-Net, denoiser, sampler, scheduler
- classifier-free guidance, guidance scale, negative prompt, ControlNet, IP-Adapter, LoRA style adapter, inpainting, outpainting, image editing, super-resolution, upscaler, latent, VAE
- text-to-image model, text-to-video model, image-to-video model, video foundation model, media foundation model, storyboard model, controllable generation, temporal upsampling
- neural codec, audio language model, music generation model, voice generation model, voice cloning model, codec language model
Reasoning Và Search
- reasoning model, thinking model, chain-of-thought training, hidden chain of thought, scratchpad, verifier-guided search, best-of-n, majority vote, self-consistency, Monte Carlo tree search, beam search over thoughts
- math reasoning, code reasoning, scientific reasoning, theorem proving, Lean, proof assistant, formal verification, symbolic regression, program synthesis, tool-augmented reasoning
- test-time compute, parallel thinking, deep research, browsing, citation search, source triangulation, evidence synthesis, answer verification, self-correction, uncertainty estimation
- RLVR, GRPO, cold-start reasoning data, reasoning trace distillation, process reward model, outcome reward model, deliberative alignment
World Models, Robotics, Embodiment
- world model, latent world model, predictive model, action-conditioned model, video world model, JEPA, I-JEPA, V-JEPA, V-JEPA 2, V-JEPA 2-AC, Genie, MuZero, Dreamer
- embodied AI, robotics foundation model, robot policy, imitation learning, behavior cloning, reinforcement learning, sim-to-real, real-to-sim, digital twin, teleoperation, DROID, Open X-Embodiment, RT-1, RT-2, AutoRT
- affordance, grasping, manipulation, navigation, motion planning, trajectory optimization, proprioception, tactile sensing, visuomotor policy, action chunking, diffusion policy
- universal AI assistant, multimodal memory, visual interpreter, prototype glasses, live camera assistant
SLM, Local, Personal AI
- Small Language Model, task-specialized model, local assistant, private assistant, personal knowledge base, offline inference, mobile inference, NPU, Apple Neural Engine, Windows Copilot Runtime
- Phi, Qwen small models, Llama small models, Mistral small models, distilled model, quantized model, edge deployment, browser inference, WebGPU, WebNN
- local RAG, personal memory, local embeddings, private vector store, encrypted memory, data residency, bring-your-own-key, federated learning, secure enclave
Neuromorphic Và Alternative Compute
- Spiking Neural Network, spike, membrane potential, leaky integrate-and-fire, event-based vision, neuromorphic computing, Loihi, TrueNorth, DVS camera
- analog AI, in-memory compute, photonic AI, wafer-scale engine, tensor core, systolic array, AI accelerator, inference ASIC
Internet, Media, Provenance
- Dead Internet Theory, bot traffic, synthetic media, AI slop, content farm, model collapse, data pollution, provenance fog, authenticity crisis
- watermarking, C2PA, content credentials, synthetic content label, invisible watermark, deepfake detection, liveness detection, media provenance, consent signal, opt-out, robots.txt for AI, data licensing
- search generative experience, AI Overviews, answer engine, citation UX, source laundering, AI SEO, GEO, answer optimization
Companies Và Labs
- OpenAI, Anthropic, Google DeepMind, Google AI, Meta AI, Microsoft AI, xAI, DeepSeek, Alibaba Qwen, Mistral AI, NVIDIA, Hugging Face, Cohere, AI2, Stability AI, Midjourney, Runway, Black Forest Labs, ElevenLabs, Perplexity, Apple, Amazon, Tencent, Baidu, ByteDance, Moonshot AI, Zhipu AI, MiniMax, Sakana AI, Adept, Scale AI, Together AI, Fireworks AI, Groq, Cerebras, Databricks Mosaic, Character.AI
- ARC Prize, Stanford HAI, EleutherAI, LAION, MLCommons, Linux Foundation
Landmark Models
- OpenAI: GPT-2, GPT-3, InstructGPT, ChatGPT, GPT-4, GPT-4o, o1, o3, GPT-5, GPT-6 Astra, DALL-E 2, DALL-E 3, Sora, Whisper
- Anthropic: Claude 3 Opus, Claude 3.5 Sonnet, Claude 4, Claude Fable 5.1
- Google DeepMind: Gemini 1.5 Pro, Gemini 2.5 Pro, Gemini 2.5 Deep Think, Gemini 3 Pro, Project Astra, Veo, Imagen, AlphaFold 2, AlphaFold 3, Genie 3, Gemma
- Meta AI: Llama 2, Llama 3, Llama 4, Segment Anything
- DeepSeek: DeepSeek-V3, DeepSeek-R1-Zero, DeepSeek-R1
- Alibaba Qwen: Qwen2.5, Qwen3, Qwen3-Omni
- Mistral AI: Mistral 7B, Mixtral 8x7B, Mistral Large
- xAI: Grok-1, Grok-3, Grok-4
- Image/video/audio: Stable Diffusion, Stable Diffusion XL, Flux, Midjourney v5, Midjourney v6, Runway Gen-2, Runway Gen-3 Alpha, Pika, Kling, HunyuanVideo, Suno, Udio
Seed Từ Backup
12 note AI cũ trong backup/*.md.bak đã được dùng làm seed và phân bổ vào các cluster bên trên.
References
- OpenAI: GPT-6 Astra, GPT-6 Astra API model page, GPT-6 Astra safety overview, GPT-5 System Card, GPT-4o System Card, o1 System Card, Sora 2, DALL-E 3, Function calling, Structured Outputs, Agents SDK tools
- Anthropic and protocols: Claude Fable, Claude Fable 5.1 and Mythos 5.1, Anthropic system cards, Model Context Protocol announcement, MCP 2026-07-28 spec announcement
- Google DeepMind and Google: Project Astra, Gemini 3 announcement, Google DeepMind model cards, Gemini 2.5 report, Veo, Genie 3, AlphaFold 3, Agent2Agent codelab
- Meta AI: Llama 4, Llama 3, V-JEPA, V-JEPA 2, Multi-token prediction paper
- DeepSeek: DeepSeek-R1 API release, DeepSeek-R1 paper, DeepSeek-V3 paper, DeepSeek-V3 GitHub, DeepSeek Transparency Center
- Other labs and infra: Qwen3 announcement, Qwen3 Technical Report, Qwen3-Omni Technical Report, Mixtral of Experts, Mistral Large, xAI Grok 4, xAI model docs, NVIDIA GB200 NVL72
- Core papers: Attention Is All You Need, GPT-2 paper, Scaling Laws for Neural Language Models, FlashAttention, Grouped-Query Attention, LoRA, QLoRA, DPO, Constitutional AI, DDPM, DDIM, Latent Diffusion, Classifier-Free Guidance, Diffusion Transformers
- Benchmarks and safety: SWE-bench, SWE-bench leaderboards, ARC-AGI-3, ARC-AGI benchmarking, OSWorld, MMLU paper, GPQA paper, HumanEval paper
- Seed notes:
backup/Transformers.md.bak,backup/Diffusion.md.bak,backup/Emergence.md.bak,backup/Harness Engineering.md.bak,backup/Jailbreaking LLMs.md.bak,backup/Multi-Token Prediction.md.bak,backup/On Having An Agent.md.bak,backup/Small LLMs — Use Cases and Limits.md.bak,backup/Spiking Neural Networks.md.bak,backup/Three Layers — Tool, MCP, Skill.md.bak,backup/World Model V-JEPA 2.md.bak,backup/Dead Internet Theory.md.bak