LLM ๊ธฐ์ด๋ถํฐ RAG, tool/MCP, agent runtime, orchestration, eval, safety, ops๊น์ง ์ชผ๊ฐฐ๋ค. ์นด๋๋ฅผ ๋๋ฅด๋ฉด ํ๊ตญ์ด ์ค๋ช , raw English, ํ๋ฆ๋, ์ฝ๋ ์ค์ผ์น, self-attack ์ง๋ฌธ์ด ์ด๋ฆฐ๋ค.
No unverified AI output lands. AI๊ฐ ๋ง๋ ์ฐ์ถ๋ฌผ์ run ledger, trace, verifier, judge, gate๋ฅผ ์ง๋์ผ ์ ๋ฌด ๊ฒฐ๊ณผ๋ฌผ์ด ๋๋ค.
๋จผ์ agent๋ฅผ ๋ง์ด ๋ง๋๋ ๊ฒ ์๋๋ค. contract, run ledger, verification loop๊ฐ ๋จผ์ ๋ค.
Where the LLM's job ends
๋๊ท๋ชจ ์ธ์ด ๋ชจ๋ธ์ ์ ํํ ๊ฒฝ๊ณ
An LLM is a text generation engine, not the whole product.
Smallest unit the model reads and writes
model์ด ์ฝ๊ณ ์ฐ๋ ์์ ๋จ์
A token is the unit the model reads and generates.
It just predicts the next token
๋ชจ๋ธ ์ถ๋ ฅ์ ๊ฐ์ฅ ๋ฎ์ ๋ ๋ฒจ ๊ฐ๊ฐ
The model generates one token at a time from a probability distribution.
Frozen parameters baked in at training
ํ๋ จ์ผ๋ก ๋ง๋ค์ด์ง ๋ด๋ถ ํ๋ผ๋ฏธํฐ
Model weights are learned parameters, not runtime memory.
Where general language ability comes from
์ผ๋ฐ ์ธ์ด/ํจํด ๋ฅ๋ ฅ์ ๋ง๋๋ ๋จ๊ณ
Pretraining gives broad capability, not your private current facts.
Teaching it to follow instructions
์ง์๋ฅผ ๋ฐ๋ฅด๋ ๋ชจ๋ธ๋ก ๋ง์ถ๋ ๋จ๊ณ
Instruction tuning makes the model better at following instructions, not perfectly reliable.
Tuning toward answers people prefer
์ข์ ๋ต๋ณ ์ ํธ๋๋ฅผ ๋ง์ถ๋ ๋จ๊ณ
Preference tuning shapes behavior, but it is not domain verification.
How the model picks each token
ํ๋ฅ ๋ถํฌ์์ ํ ํฐ์ ๊ณ ๋ฅด๋ ๋ฐฉ์
Sampling controls how the model chooses among possible next tokens.
Tuning output variety vs. stability
์ถ๋ ฅ ๋ค์์ฑ๊ณผ ์์ ์ฑ ์กฐ์
Temperature and top-p trade off variety against stability.
The model's working memory for one request
ํ ์์ฒญ์์ ๋ชจ๋ธ์ด ๋ณผ ์ ์๋ ์์ ๊ธฐ์ต
The context window is the model's working memory for one request.
Turning text into searchable vectors
ํ ์คํธ๋ฅผ ๊ฒ์ ๊ฐ๋ฅํ ๋ฒกํฐ๋ก ๋ฐ๊ฟ
An embedding turns text into a vector for semantic search.
Trading off quality, cost, speed, and tool use
ํ์ง, ๋น์ฉ, ์๋, tool ๋ฅ๋ ฅ์ trade-off
Model selection is a product and systems trade-off.
Confident answers with nothing backing them
๊ทธ๋ด๋ฏํ์ง๋ง ๊ทผ๊ฑฐ ์๋ ์ถ๋ ฅ
A hallucination is plausible output without sufficient grounding.
The system / developer / user instruction layers
system, developer, user instruction์ ๊ณ์ธต
Message roles separate product policy from the user's current request.
Locking down goal, input, output, and the don'ts
๋ชฉํ, ์ ๋ ฅ, ์ถ๋ ฅ, ๊ธ์ง์ฌํญ์ ๊ณ ์
A prompt contract defines task, inputs, output shape, constraints, and failure behavior.
Pinning the pattern down with examples
์ํ๋ ํจํด์ ์์๋ก ๊ณ ์
Few-shot examples teach the model the pattern you want.
Machine-readable results instead of free text
free text ๋์ machine-readable ๊ฒฐ๊ณผ
Structured output turns model text into something the system can validate and route.
Spend reasoning effort where it pays off
๋์ด๋์ ๋ง๊ฒ ์๊ฐ ๋น์ฉ์ ๋ฐฐ์
Reasoning depth should match task complexity and risk.
Pulling in evidence before answering
retrieval๋ก ์ฆ๊ฑฐ๋ฅผ ๋ฃ๊ณ ๋ตํ๊ฒ ํจ
RAG gives the model grounded evidence; it does not make it automatically truthful.
Splitting docs into retrievable pieces
๋ฌธ์๋ฅผ ๊ฒ์ ๊ฐ๋ฅํ ์กฐ๊ฐ์ผ๋ก ๋๋
Chunking controls the unit of retrieval.
Brute-force kNN over millions of embeddings is too slow โ ANN indexes trade recall against latency and memory
์๋ฐฑ๋ง ์๋ฒ ๋ฉ์์ ์์ ํ์ kNN์ ๋๋ฆฌ๋ค โ ANN ์ธ๋ฑ์ค๋ก recallยท์ง์ฐยท๋ฉ๋ชจ๋ฆฌ๋ฅผ ์ ์ถฉํ๋ค
Brute-force kNN over millions of vectors is O(N times d), so we accept slightly lower recall and use an ANN index โ HNSW for speed and high recall at higher memory cost, IVF for lower memory with a tunable probes count.
Narrowing the search space first
๊ฒ์ ์ ์ ๋ฒ์๋ฅผ ์ขํ
Metadata filters keep retrieval scoped before similarity search.
Combining keyword and vector search
keyword + vector ๊ฒ์ ๊ฒฐํฉ
Hybrid search combines lexical matching with semantic similarity.
Re-sorting candidates by real relevance
ํ๋ณด ๋ฌธ์์ ์ค์ ๊ด๋ จ๋๋ฅผ ๋ค์ ์ ๋ ฌ
Reranking improves the order and relevance of retrieved evidence.
Tying answers back to evidence
๋ต์ ์ฆ๊ฑฐ์ ๋ฌถ์
Grounding ties the answer to evidence the system can inspect.
Squeezing long context down to what the task needs
๊ธด ์ ๋ณด๋ฅผ ์์ ์ ๋ง๊ฒ ์์ถ
Context compression saves space but can lose evidence.
Session vs. user vs. semantic vs. episodic memory
session, user, semantic, episodic memory ๊ตฌ๋ถ
Memory is not one bucket; different memories have different lifetimes and risks.
Stale evidence poisoning the answer
์ค๋๋ ์ฆ๊ฑฐ๊ฐ ๋ต์ ๋ง์นจ
Stale context makes the model confidently use outdated facts.
Model requests a structured action
๋ชจ๋ธ์ด ๊ตฌ์กฐํ๋ ํ๋์ ์์ฒญํจ
A tool call is a structured action request, not permission to do anything.
The tool's input contract
tool ์ ๋ ฅ ๊ณ์ฝ
A function schema is the contract for a tool call.
Don't trust the model's tool args
๋ชจ๋ธ์ด ๋ธ tool args๋ฅผ ๋ฏฟ์ง ์์
Never execute tool arguments just because the model produced them.
Tools that don't double-fire on retry
์ฌ์๋ํด๋ ์ค๋ณต ์คํ๋์ง ์๋ ๋๊ตฌ
A retried tool call must not create duplicate side effects.
Feeding tool failures back into agent state
tool ์คํจ๋ฅผ agent state๋ก ๋๋๋ฆผ
Tool failures should become typed observations, not chaos.
Isolating code, file, and command execution
์ฝ๋/ํ์ผ/๋ช ๋ น ์คํ ๊ฒฉ๋ฆฌ
A sandbox limits the blast radius of tool execution.
Human approval before risky tool calls
์ํํ tool call ์ ์ฌ๋ ์น์ธ
High-risk actions need approval before execution, not after damage.
Standard way to expose external tools
์ธ๋ถ ๋๊ตฌ๋ฅผ ํ์ค ๋ฐฉ์์ผ๋ก ๋ ธ์ถ
MCP is an open protocol for exposing external tools and data to agents in a standard way; permission, schema validation, and audit live in my own gateway on top of it.
Runtime that loops the model plus tools toward a goal
๋ชฉํ๋ฅผ ํฅํด ๋ชจ๋ธ+๋๊ตฌ๋ฅผ ๋ฐ๋ณต ์คํํ๋ runtime
An agent is a controlled loop over reasoning, tools, and state.
Reason, act, observe, repeat
reason + act + observe ๋ฐ๋ณต
ReAct means reason, act, observe, and repeat under control.
Task state carried through the agent loop
agent loop๊ฐ ๋ค๊ณ ๊ฐ๋ ์์ ์ํ
Agent state is the memory of the current run.
Current state of a conversation or session
๋ํ/์์ ์ธ์ ์ ํ์ฌ ์ํ
Session state is temporary working state, not permanent memory.
When the agent decides to stop
agent๊ฐ ์ธ์ ๋ฉ์ถ๋์ง
Every agent loop needs explicit stop conditions.
How failures get retried
์คํจ๋ฅผ ์ด๋ป๊ฒ ๋ค์ ์๋ํ ์ง
Retries need policy, not hope.
Next path when the best one fails
์ฝํ ๊ฒฝ๋ก ์คํจ ์ ๋ค์ ๊ฒฝ๋ก
A reliable agent has a fallback path when confidence is low.
Central control over the workflow and its gates
workflow์ gate๋ฅผ ํต์ ํ๋ ์ค์ ์ ์ด
Agents do not freely chat; the orchestrator controls the workflow.
Structured output instead of free-form text
free-form ๋์ ๊ตฌ์กฐํ๋ ์ถ๋ ฅ
A multi-agent system needs contracts, not free-form reports.
One object multiple agents work on
์ฌ๋ฌ agent๊ฐ ๊ฐ์ ๋์๋ฌผ์ ๊ฒํ
Multi-agent workflows need a shared, versioned artifact.
Every AI job recorded as a run
๋ชจ๋ AI ์์ ์ run์ผ๋ก ๊ธฐ๋ก
No run ledger, no trustworthy AI workflow.
draft โ verified โ accepted/fix/rejected
AI output starts as a draft, not as accepted work.
Checking with multiple independent criteria at once
์๋ก ๋ค๋ฅธ ๊ธฐ์ค์ผ๋ก ๋์์ ๊ฒ์ฆ
Parallel verifiers inspect the same artifact through different lenses.
Domain-specific verification runner
domain๋ณ ๊ฒ์ฆ์ ์คํ
No unverified AI output lands.
Merges multiple verifier results into a final call
์ฌ๋ฌ verifier ๊ฒฐ๊ณผ๋ฅผ ํฉ์ณ ์ต์ข ํ์
The judge merges findings into one verdict.
Measuring AI behavior reproducibly
AI behavior๋ฅผ ์ฌํ ๊ฐ๋ฅํ๊ฒ ์ธก์
An eval dataset makes AI quality measurable across versions.
The must-never-miss representative cases
์ ๋ ๋์น๋ฉด ์ ๋๋ ๋ํ ์ผ์ด์ค
A golden set is the small set of cases that must never regress.
Trading off false positives against missed issues
false positive์ missed issue์ ๊ท ํ
Precision asks how many findings are real; recall asks how many real issues were found.
Turn fixed failures into tests so they don't come back
๊ณ ์น ์คํจ๋ฅผ ๋ค์ ์ ๋์น๊ฒ ํ ์คํธํ
A rejected or fixed run should become an eval case.
External text hijacking the agent's instructions
์ธ๋ถ ํ ์คํธ๊ฐ agent ์ง์๋ฅผ ํ์ทจ
Prompt injection is untrusted content trying to control the model.
The agent leaking sensitive data out
agent๊ฐ ๋ฏผ๊ฐ ์ ๋ณด๋ฅผ ๋ฐ์ผ๋ก ํ๋ฆผ
Data exfiltration is the agent leaking data it should not reveal.
Giving the agent only the access it needs
agent์๊ฒ ํ์ํ ๊ถํ๋ง ์ค
An agent should only get the tools and data it needs for this task.
Controlling how sensitive data is collected, stored, and exposed
๋ฏผ๊ฐ ์ ๋ณด์ ์์ง, ๋ณด๊ด, ๋ ธ์ถ ํต์
Sensitive data must be minimized, redacted, and checked before output.
Risky outputs need human sign-off
์ํํ ์ฐ์ถ๋ฌผ์ ์ฌ๋์ด ์น์ธ
Automation needs gates where risk is high.
Logs tool calls, latency, cost, and failure reasons
tool call, latency, cost, failure reason ๊ธฐ๋ก
If I cannot trace it, I cannot operate it.
Keeping token, tool, and retry costs in check
ํ ํฐ, tool, retry ๋น์ฉ ํต์
AI cost is an operational metric, not an afterthought.
Designing within the time a user is willing to wait
์ฌ์ฉ์๊ฐ ๊ธฐ๋ค๋ฆด ์ ์๋ ์๊ฐ ์์ ์ค๊ณ
Latency budget decides what runs synchronously and what becomes async.
Preventing abuse and runaway cost
๋จ์ฉ๊ณผ ๋น์ฉ ํญ๋ฐ ๋ฐฉ์ง
Rate limits protect cost, reliability, and external APIs.
Shipping a new prompt, model, or agent to a slice of traffic first
์ prompt/model/agent๋ฅผ ์ผ๋ถ ํธ๋ํฝ์๋ง ์ ์ฉ
Canary rollout limits the blast radius of an AI behavior change.
Watching for quality changes over time
์๊ฐ์ด ์ง๋๋ฉฐ ํ์ง์ด ๋ณํ๋์ง ๊ฐ์
Drift monitoring catches quality changes after deployment.
Decide the value before touching the tech
๊ธฐ์ ์ ์ ๋ง๋ค ๊ฐ์น๋ถํฐ ํ์
Technical feasibility does not mean we should build it.