Token
Smallest unit the model reads and writes
at scale์์ ๋๊ฐ ์๋๋ผ ํฐ ํธ๋ํฝ/๋ฐ์ดํฐ์์.phrase
์ง์ญ: ๊ท๋ชจ๊ฐ ์ปค์ก์ ๋
๋์์ค: ๋ง์ผํ ์ฒ๋ผ ์ฐ์ง ๋ง๊ณ ์ด๋ค ์์น์์ ๊นจ์ง๋์ง ๋ถ์ฌ์ผ ํ๋ค.
This works locally but breaks at scale.
chunk๊ฒ์๊ณผ RAG๋ฅผ ์ํด ๋ฌธ์๋ฅผ ๋๋ ๋จ์.term
์ง์ญ: ์กฐ๊ฐ
๋์์ค: ๋๋ฌด ์์ผ๋ฉด ๋ฌธ๋งฅ ๋ถ์กฑ, ๋๋ฌด ํฌ๋ฉด noise๊ฐ ๋์ด๋๋ค.
The answer cites the retrieved chunk.
context windowํ ์์ฒญ์์ ๋ชจ๋ธ์ด ๋ณผ ์ ์๋ ์ต๋ ํ ํฐ ๋ฒ์.term
์ง์ญ: ๋ฌธ๋งฅ ์ฐฝ
๋์์ค: ๋ชจ๋ธ์ ์ฅ๊ธฐ๊ธฐ์ต์ด ์๋๋ผ ๊ทธ ์๊ฐ์ ์์ ๊ณต๊ฐ์ด๋ค.
Older turns fall out of the context window.
fall outcontext๋ cache ๋ฒ์์์ ๋น ์ ธ์ ๋ ์ด์ ๋ณด์ด์ง ์๊ฒ ๋๋ค.verb
์ง์ญ: ๋ฐ์ผ๋ก ๋จ์ด์ง๋ค
๋์์ค: LLM history, cache, sliding window์์ ์์ฐ์ค๋ฝ๊ฒ ์ด๋ค.
Old messages fall out of the context window.
latency์์ฒญ ํ๋๊ฐ ๋๋๊ธฐ๊น์ง ๊ฑธ๋ฆฌ๋ ์๊ฐ.term
์ง์ญ: ์ง์ฐ์๊ฐ
๋์์ค: ํ๊ท ๋ณด๋ค percentile์ด ์ค์ํ๋ค.
Latency is measured per request.