Skip to main content

Command Palette

Search for a command to run...

UNPACKING LLM,s

Updated
•9 min read•View as Markdown

Hey , I am karan and today i am going to unpack the magical black box of llm, you will today learn about how gpt works behind the scene .

lets first understand what is llm?

LLM stands for Large language model.

  • large → trained on massive amounts of text data (billions to trillions of words)

  • Language → it works with human language (text) — understanding it, generating it, reasoning over it

  • Model → it's a neural network (typically Transformer-based, as we discussed) trained to predict and generate text

Problem that llm solves:

LLMs solve the problem of contextual language understanding. Earlier NLP systems could process text but struggled to understand the meaning of words in different contexts. LLMs learn language patterns from massive datasets, allowing them to understand context and generate coherent, human-like responses.

Model Made by
GPT-4 / GPT-4o / GPT-5 OpenAI
Claude (Opus, Sonnet, Haiku) Anthropic
Gemini Google DeepMind
Llama (2, 3, 4) Meta
Mistral / Mixtral Mistral AI
Grok xAI
DeepSeek DeepSeek

Common Applications in Daily Life

  • Chatbots & virtual assistants — ChatGPT, Claude, Siri/Alexa-style assistants for Q&A and help with tasks

  • Writing assistance — drafting emails, essays, resumes, blog posts; grammar/style correction (e.g., Grammarly)

  • Coding help — autocompletion, debugging, generating code (GitHub Copilot, Claude Code, Cursor)

  • Search & summarization — summarizing long articles, documents, meeting notes

  • Translation — real-time language translation (Google Translate increasingly uses LLM-style models)

  • Customer support — automated support chatbots that handle queries 24/7

  1. What Happens When You Send a Message to ChatGPT?

Typing a prompt

It starts simply enough — you type a sentence, a question, a request. To you, this is just plain English (or whatever language you're using). To the computer, it's meaningless until it's been transformed.

Processing your message

Once you hit send, your text doesn't go straight into some kind of search engine or database lookup. Instead, it goes through a pipeline:

Your text is broken into small chunks called tokens Each token is converted into a number, then into a vector (a list of numbers) that captures meaning Position information is added, so the model knows the order your words appeared in The whole sequence passes through many layers of a neural network called a Transformer, which builds up a deep, contextual understanding of what you're asking

By the end of this, the model doesn't have "your sentence" anymore — it has a rich mathematical representation of the meaning of your sentence.

Generating a response

Now comes the generation part. The model doesn't write out a full sentence in one shot. It predicts one token at a time:

It looks at everything so far (your prompt + whatever it has generated already) It calculates a probability for every possible next token in its vocabulary It picks the most likely one (or samples from the top candidates, depending on settings) That token gets added to the sequence, and the whole process repeats This continues until the model produces a special "end" signal or hits a length limit

So the reply you see being "typed out" isn't a performance for your benefit — that's genuinely how it's being built, one predicted piece at a time.

Why responses are not copied from the internet

This is one of the most common misconceptions. ChatGPT isn't searching a database of pre-written answers and pasting one back to you. During training, the model read enormous amounts of text and adjusted billions of internal weights to get better at predicting the next token. It's not storing sentences — it's storing patterns, relationships, and statistical structure in its weights.

When it responds to you, it's generating new text token by token, based on probabilities learned from patterns in that training data — not retrieving and copying a stored document. That's why it can answer questions that were never asked anywhere before, and why it can be wrong: it's predicting plausible continuations, not looking up facts.

  1. Why Computers Don't Understand Human Language

Text vs numbers

Here's the core issue: computers, at the deepest level, only understand numbers. Every operation a processor performs — addition, comparison, matrix multiplication — works on numerical values. Words like "cat," "run," or "happiness" mean nothing to a CPU or GPU; they're just squiggly shapes to us that carry meaning, but to a machine they're not even data yet.

Why computers need everything converted into numbers

For a neural network to process language, every single piece of text has to first be turned into numbers it can do math on. This isn't optional — it's a hard requirement of how neural networks function. Weights, activations, gradients — all of it is numerical.

So before any "understanding" can happen, we need a reliable way to turn text into numbers, and back again. That's where tokens come in.

Introduction to tokens

A token is the basic unit the model actually operates on — not necessarily a whole word, sometimes a piece of a word, a whole word, or even punctuation. Every token has a corresponding number (an ID) in the model's vocabulary. This numeric ID is the first step in the transformation from human language into something a neural network can process.

  1. Tokenization

What tokens are

Tokenization is the process of breaking text into these smaller units — tokens — and mapping each one to a unique number based on a predefined vocabulary the model was trained with.

Why tokenization is needed

Language is incredibly varied — new words, slang, typos, made-up words, different languages. If a model only understood whole words, it would constantly hit words it had never seen before and have no way to represent them. Tokenization solves this by breaking text into smaller, reusable pieces, so that even unfamiliar words can usually be built from familiar sub-word pieces.

It also keeps the vocabulary a manageable size. Instead of needing a separate entry for every possible word in every possible form ("run," "running," "runner," "runs"...), the model can share common sub-word pieces across many words.

Words vs tokens

A common misconception is that 1 word = 1 token. That's often not true:

Common short words are usually a single token Longer or less common words often get split into multiple tokens Punctuation and spaces can be their own tokens

This is why LLM usage is often measured in "tokens" rather than "words" — a 100-word paragraph might be 130 tokens, or more, depending on the vocabulary used.

Simple examples

Take the sentence: "Tokenization is fascinating"

A tokenizer might break it down like this:

"Tokenization" → "Token" + "ization" (2 tokens) "is" → "is" (1 token) "fascinating" → "fascinat" + "ing" (2 tokens)

So a 3-word sentence could become 5 tokens. Each of these tokens then gets mapped to a number:

"Token" → 4521 "ization" → 872 "is" → 15 "fascinat" → 9034 "ing" → 202

Those numbers are what actually flow into the model — this is the true starting point of everything the network does next (embedding, positional encoding, attention, and so on, as we discussed earlier).

  1. Transformers

What a Transformer is

The Transformer is the neural network architecture that powers nearly every modern LLM. It was introduced in 2017 in a paper titled "Attention Is All You Need." At its core, a Transformer takes a sequence of tokens (as numbers/vectors) and processes them using a mechanism called self-attention, which lets every token look at every other token in the sequence and figure out how relevant they are to each other.

Why it changed AI

Before Transformers, models like RNNs and LSTMs processed language sequentially — one word at a time, in order, carrying forward a "memory" of what came before. This had two big problems:

It was slow — sequential processing couldn't be parallelized, so training on huge datasets took a very long time It struggled with long-range context — by the time an RNN reached the end of a long sentence or paragraph, it had often "forgotten" details from the beginning

Transformers solved both problems at once. Because self-attention lets every token connect directly to every other token — regardless of distance — the model doesn't lose track of something said much earlier. And because there's no step-by-step sequential dependency during training, the whole sequence can be processed in parallel, making it possible to train on internet-scale datasets in a practical amount of time.

How it helps understand language

Self-attention gives the model the ability to weigh relationships between words dynamically. Consider: "The trophy didn't fit in the suitcase because it was too big."

What does "it" refer to — the trophy or the suitcase? A human resolves this instantly from context. Self-attention lets the model do something similar: when processing the word "it," attention scores let the model weigh how relevant "trophy" and "suitcase" each are, and the surrounding context ("too big") tips the scales toward "trophy."

Stack many of these self-attention layers together, interleaved with feed-forward layers (as we covered earlier), and the model builds up increasingly rich, contextual understanding of the entire input — not just what each word means, but what it means in this specific sentence.

Why almost every modern LLM uses Transformers

The combination of parallelizable training and strong long-range context handling made Transformers dramatically more effective and efficient to scale than anything before them. Bigger models trained on more data kept getting better, seemingly without hitting the same walls older architectures did.

This is why virtually every major LLM today — GPT, Claude, Gemini, Llama, Mistral, and others — is built on the Transformer architecture (typically using just the decoder half, as we touched on earlier, for pure text-generation tasks). It became the foundation the entire field converged on, and to date nothing has meaningfully displaced i