Part 1: How LLMs work - tokens, vectors, attention
Module 10 · Sat 19 Sep
Hinglish mein suno
LLMs koi magic chat box nahi hain. Yeh token-based systems hain jinka behaviour context, prompting aur product design par depend karta hai - isliye PMs ko basics pata hone chahiye.
- Tokens cost aur processing ki real unit hain. Token count ≠ word count - ek word kai tokens mein toot sakta hai.
- Vectors woh numbers hain jo meaning represent karte hain; zyada dimensions = richer meaning. Attention pehle ke important tokens ko weight deta hai, aur plain vector ko contextual vector bana deta hai.
- Model next-token prediction se likhta hai, ek-ek token karke.
Product levers: temperature (support, enterprise ya factual assistants ke liye low; creative output ke liye high), max tokens (har task ke hisaab se set karo, default se nahi), input limits aur model choice (bada = zyada capable, zyada mehenga). Cost, latency aur user need ko saath mein socho.
Zyada context hamesha better nahi - relevant context help karta hai, noisy context nuksaan karta hai. Bade documents paste mat karo; relevant parts retrieve karo (class ne ise RAG aur retrieval systems se joda). Domain context hi business AI ko useful banata hai, sirf impressive nahi. Aur prototypes aasaan hain - validation mushkil part hai.