Search by name, acronym, or category. Filter by how deep it goes. Every term links to what to learn before it and after it.
No terms match that search.
What a language model actually is, and the mechanics underneath it.
Attention is the mechanism a transformer uses to decide, for each token, how much every other token in the sequence should influence it, letting the model connect a pronoun to the noun it refers to, or a closing bracket to the one it opens.
AdvancedAn LLM is a neural network trained on huge amounts of text to predict the next word in a sequence, well enough that stringing those predictions together produces coherent answers, summaries, and code.
BeginnerA token is the unit an LLM actually reads and writes, a chunk of text, often a word piece or a few characters, not a whole word or letter.
BeginnerThe transformer is the neural network architecture nearly every modern LLM is built on, defined by its use of self-attention to weigh every token against every other token at once, instead of processing text strictly left to right.
IntermediateGiving a model access to information it was never trained on.
An embedding is a list of numbers, typically hundreds or thousands of them, that represents the meaning of a piece of text (or image, or audio) as a point in space, so that similar meanings end up as nearby points.
IntermediateRAG is a pattern where, before an LLM answers a question, the system first retrieves relevant documents from an external source and includes them in the prompt, so the model answers from that specific material instead of only from what it memorized during training.
IntermediateReranking is a second, more precise scoring pass over a retrieval system's initial results, reordering them by relevance before the top few are handed to the LLM, catching cases the first-pass search got wrong.
AdvancedRetrieval is the step in RAG that finds and returns the most relevant pieces of text for a given query, either by meaning (semantic search over embeddings), by exact terms (keyword search), or both at once (hybrid search).
IntermediateA vector database stores embeddings and answers "which of these millions of vectors are closest to this one" quickly, using specialized indexing rather than comparing against every stored vector one by one.
IntermediateModels that take actions, not just answer questions.
An AI agent is an LLM wired into a loop where it can take actions, calling tools, reading the results, and deciding what to do next, rather than just producing one text response to one prompt.
BeginnerMCP is an open, standardized protocol for connecting an AI application to external tools and data sources, so a tool built once can be used by any MCP-compatible model or app, instead of every integration being built one-off.
IntermediateTool calling is a model's ability to output a structured request to run a specific function, with specific arguments, instead of (or alongside) a plain text answer, letting application code execute it and feed the result back in.
IntermediateWhat happens between a prompt and a response, and how it's made fast.
Decode is the phase of inference that generates the response one token at a time, each new token depending on every token that came before it, including the ones just generated.
AdvancedInference is the act of running a trained model to produce an output, generating a response to a prompt, as opposed to training, which is the earlier, much more expensive process of teaching the model in the first place.
BeginnerThe KV cache stores the key and value vectors a transformer computes for every token it has already processed, so generating the next token can reuse that work instead of recomputing attention over the whole sequence from scratch.
AdvancedModel serving is the infrastructure layer that takes a trained model and makes it available to answer real requests, handling batching, queueing, and GPU allocation so many requests can be served efficiently at once.
IntermediatePrefill is the first phase of inference, processing the entire input prompt in parallel to build up the KV cache, before the model generates a single token of the response.
AdvancedQuantization reduces the numeric precision used to store a model's weights, for example from 16-bit to 8-bit or 4-bit numbers, shrinking memory use and often speeding up inference, at some cost to accuracy.
AdvancedvLLM is an open-source inference engine built around paged attention, a technique that manages the KV cache in fixed-size blocks (like an operating system manages memory pages) to dramatically reduce wasted GPU memory.
InfrastructureThe hardware and routing layer underneath every request.
An AI gateway is a layer that sits between an application and the various AI model providers it uses, handling routing, rate limiting, retries, and cost tracking through one consistent interface instead of the app calling each provider directly.
IntermediateAI infrastructure is the full stack underneath an AI application, GPUs, networking between them, storage for weights and data, and the serving and orchestration software tying it together, none of which is the model itself, but all of which decides whether the model actually runs well.
InfrastructureCUDA is NVIDIA's programming platform for writing software that runs directly on its GPUs, and it's the foundation nearly every AI training and inference framework is built on top of.
InfrastructureA GPU is a processor built with thousands of small cores designed to do the same simple operation on many pieces of data at once, which happens to be exactly the kind of math (matrix multiplication) that neural networks are built from.
BeginnerA model router is the component, usually inside an AI gateway, that decides which specific model handles a given request, based on rules like cost, latency, task type, or current load.
IntermediateRunning AI systems in production: watching them, measuring them, keeping them up.
AI evaluation is the practice of systematically measuring how well an AI system performs, against a fixed set of test cases and a defined rubric, rather than trusting a handful of manual spot checks.
IntermediateAI observability is the practice of instrumenting an AI system so you can see, after the fact, exactly what happened on any given request, which prompt went in, which tools were called, how many tokens were used, how long each step took, and what the model actually output.
IntermediateAI reliability is keeping an AI system available and correct under real-world conditions, provider outages, traffic spikes, malformed inputs, by building in fallback paths and graceful degradation rather than assuming the happy path always holds.
AdvancedLLMOps is the set of practices for running LLM-based systems reliably in production, prompt versioning, evaluation pipelines, deployment, monitoring, and rollback, the AI-specific evolution of what DevOps and MLOps do for traditional software and machine learning systems.
IntermediateWhat can go wrong, and what it costs.
AI cost is what it actually costs to run an AI system in production, driven mainly by token usage (input and output, priced separately) multiplied by each model's per-token price, plus infrastructure and any added steps like retrieval or reranking.
BeginnerAI security covers the set of risks specific to LLM-based systems, prompt injection, data exfiltration through a tool call, an agent taking an action it shouldn't, that don't map cleanly onto traditional application security threats.
IntermediateGuardrails are checks placed before or after an LLM call, filtering unsafe input, catching policy violations in output, or verifying a response is actually grounded in the retrieved context, rather than relying on the model to police itself.
Intermediate← Knowledge map · Discover AI Speed →