A robot who can't taste food learns to understand it anyway. On the way you'll learn vectors, dot products and matrices: the math under every AI model. In 2D, in 3D, and in 768 dimensions.
~60 minNo math background needed10 chapters + quiz · bonus parts are optionalTrue stories from 2,000 years of math
Ravi's taste map, in 3D. Drag to turn it.
Ravi runs a small food stall near the bus stand in Vellore. Dosa in the morning, parotta in the evening, and sweets all day.
One morning a big box arrives from his nephew, who studies engineering in Chennai. Inside is a shiny robot.
Hello. I am Bolt. I can cook 200 dosas an hour. I never get tired.
Wonderful! Then tell me, which dish should I give to Priya? She comes every day.
I do not know. I cannot taste. I only understand numbers.
Ravi scratches his head. If Bolt only understands numbers, then he must turn food into numbers. This is exactly the problem every AI has. A computer cannot see a photo, read a word, or taste a dosa. It can only work with numbers. So the first job in AI is always: turn the thing into numbers.
How this page works: read the story, play with every Try it box, and answer the Pause and predict questions before peeking. Guessing first, even wrongly, is what makes ideas stick. Words with a dotted underline explain themselves when you tap them. Gold story cards, Quick refreshers and Bonus parts are closed; open them only if you want to.
Chapter 1
A dish is a list of numbers
Ravi takes a notebook. For every dish he writes two scores from 0 to 10: how spicy it is, and how sweet it is.
Dish
Spicy
Sweet
As a list
[6, 1]. I understand this! Masala Dosa is a 6 and a 1.
Key idea · Vector
A list of numbers like [6, 1] is called a vector. You can also draw it: start at the corner (0, 0), walk 6 steps right and 1 step up. The arrow you just drew is the vector. Same thing, two ways to look at it: a list for the computer, an arrow for your eyes.
The length of the arrow tells you how strong the taste is overall. We find it with the rule for right-angled triangles: length = √(6² + 1²) ≈ 6.08.
🧮Quick refresherSquares (²) and square roots (√)Show me
Squaring a number means multiplying it by itself. The small ² is how we write it: 6² = 6 × 6 = 36, and 1² = 1 × 1 = 1.
A square root (√) goes backwards: which number, times itself, gives this? √36 = 6, because 6 × 6 = 36. √37 isn't a whole number. It's a tiny bit more than 6, about 6.08.
So √(6² + 1²) = √(36 + 1) = √37 ≈ 6.08. The wavy ≈ means "about". You never have to do this by hand; the page and the computer do it for you.
Pause and predict
Pongal is [2, 1]. "Super Pongal" is [4, 2]: twice as spicy and twice as sweet. Where is it on the map?
Aha!Same direction, just twice as far. So an arrow tells you two things. Its direction tells you what kind of taste it is. Its length tells you how strong. Keep this in your pocket; it's the secret behind Chapter 4.
🧮Quick refresherReading a map with two numbersShow me
A point on a flat map needs two numbers. The first says how far to walk across (to the right). The second says how far to walk up. So [6, 1] means: start at the corner, 6 steps right, 1 step up.
The corner itself is [0, 0]: no steps at all. Order matters: [1, 6] is a different place (1 right, 6 up).
Try it · Invent a dishDrag the orange dot
To do
Bolt's challenge: invent a dish that is very spicy and very sweet (7 or more on both).
One more number: going 3D
Bolt, spicy and sweet are not enough. Lemon rice and tamarind rice are sour! Let me add a third score.
Now each dish is 3 numbers: [spicy, sweet, sour]. Two numbers needed a flat page. Three numbers need a room: across, back, and up. Each separate kind of score is called a dimension. Ravi's map just went from 2 dimensions to 3.
Try it · The taste room 3DDrag the picture to turn it
To do
Bolt's challenge: turn the room so you look straight down from the ceiling. Watch what happens to the sour direction.
The length rule still works in 3D. You just add one more square: length = √(spicy² + sweet² + sour²). Remember that. It's about to take us somewhere strange.
You just learned
A vector is a list of numbers, like [6, 1]. You can also draw it as an arrow.
Each kind of score (spicy, sweet, sour) is one dimension.
An arrow's direction says what kind of thing it is. Its length says how strong.
Chapter 2
Beyond three dimensions
Ravi gets excited. Salty! Crunchy! Oily! Hot or cold! Price! Soon every dish has 8 numbers.
Ravi, I cannot draw this. A room only has 3 directions at right angles to each other. Where does the 4th arrow go?
Nobody can picture it, Bolt. Not me, not the smartest mathematician alive. But do you need a picture to do the math?
Here's the trick: forget arrows and draw each number as a bar. Bars work for 2 numbers, 8 numbers, or 768 numbers. And the length rule doesn't care how many numbers there are. You keep adding squares.
Try it · Add more dimensionsSlide to add scores
You can picture it as
Pani Puri length
Distance between the two dishes
Aha!
The math works the same in every dimension, even ones you can't see. You can't picture 8 dimensions, but you can compute in 8 dimensions without breaking a sweat. That's exactly how AI works: it computes in hundreds or thousands of dimensions that no human can picture.
★Bonus · optionalCan we get a feel for 4D?Open bonus
A mind-bending detour. Skip it if you like; nothing later needs it.
A little, using shadows. Hold a wire cube under a lamp and its shadow on the table is a flat 2D drawing: a square inside a square, joined at the corners. A 2D creature living on the table would only ever see that shadow. In the same way, we can draw the 3D shadow of a 4D cube. It looks like a cube inside a cube.
Pause and predict
A square has 4 corners. A cube has 8. How many corners does a 4D cube have?
Aha!16. Every new dimension doubles the corners: take the shape, make a copy, push the copy out in the new direction, and join each corner to its twin. So 2, 4, 8, 16, 32… A cube in 768 dimensions has 2768 corners. That's a number with 232 digits, far more than the number of atoms in the universe (about 1080). High dimensions are roomy, and that's why AI can fit so much meaning in there.
Try it · The shadow of a 4D cube 4DDrag to turn · slider turns through the 4th direction
Corners
Edges
Corner as numbers
Orange edges are the ones that point along the newest direction. In the 4D cube, the small inner cube is the part "farther away" in the 4th direction, so its shadow is smaller.
To do
Bolt's challenge: pick the 4D cube and turn it a full quarter turn (90°) through the 4th direction. Watch the inner and outer cubes swap places.
From dishes to words
My nephew says AI chatbots do the same thing with words.
But a word has no spice or sugar. What numbers would I give it?
Good question. Imagine describing the word king with scores. Is it a person? (yes: 9) Is it royal? (10) Is it food? (0) Is it old-fashioned? (7) To capture a word's meaning you'd need lots of scores, so lots of dimensions.
Here's the surprise: nobody fills in these scores by hand. A computer reads billions of sentences. Every word starts with random numbers. Each time two words show up in similar sentences ("the king ruled the land", "the queen ruled the land"), the computer nudges their numbers a little closer. After enough reading, words with similar meanings end up with similar lists. The computer's scores don't have neat names like "royal". They're just useful.
New wordEmbedding: the list of numbers a computer has learned for a word (or a photo, a song, a shopper). It sounds fancy, but it only means "this thing, written as a vector".
Why 768 numbers? There's no magic in it. It's the size the engineers picked for Google's BERT model in 2018: enough numbers to hold a lot of meaning, few enough to run fast. (It also splits neatly into 12 groups of 64, which their design needed.) Bigger models use longer lists. GPT-3 (2020) uses 12,288 numbers for every word, or more exactly for every piece of a word.
Try it · What a word looks like to AIPick a size
Each little cell is one number: blue for positive, red for negative, darker for bigger. These are made-up numbers to show the idea, not a real model's. Notice that king and queen look alike and dosa doesn't. Your eyes can't read 768 numbers, but Chapter 4 teaches the one calculation that compares them in a blink. Sneak peek:
To do
Bolt's challenge: press 2 numbers, then 768 numbers. With which size do king and queen clearly look alike?
You just learned
You can't picture more than 3 dimensions, but the math works the same with 8 or 768 numbers.
An embedding is the list of numbers a computer learns for a word (or photo, or song).
Similar words end up with similar lists. Nobody fills them in by hand; the computer learns them by reading.
Interlude
Who invented the arrow?
Ravi, who invented vectors?
Hmm. I don't know, Bolt. Somebody clever, long ago?
Many somebodies, over 2,000 years. It took rice farmers, a bridge, a forgotten schoolteacher and a nasty fight. Here's the short version.
~200 BCE–100 CEChina
Grids of numbers.The Nine Chapters on the Mathematical Art solves problems about bundles of rice by laying numbers out in a grid on a counting board and cancelling them out step by step. That grid is a matrix, about 1,600–2,000 years before Europe gave it a name.
1637France
Points become numbers. Descartes (and Fermat) name every point with numbers.
1687England
Adding pushes. Isaac Newton writes that a body pushed by two forces at once moves along the diagonal of a parallelogram. That's adding arrows (Chapter 3).
1797Denmark
The ignored surveyor. Caspar Wessel, a map surveyor, shows how to add and turn arrows on paper. Almost nobody reads it for about 100 years.
1843Dublin
The bridge. William Rowan Hamilton carves a formula into a bridge (story below). He gives us the words scalar and, in its modern sense, vector.
1844Germany
Any number of dimensions. Grassmann's book that nobody read.
1880sUSA & England
Modern vectors. Josiah Willard Gibbs at Yale and Oliver Heaviside in England, separately, build the simple vector toolkit we still use today. Heaviside rewrites Maxwell's 20 equations of electricity and magnetism as 4 short vector equations.
Chapter 3
Mixing and doubling
Ravi anna, can I get Pani Puri with a little Kesari on top? I want to try something new.
What will that taste like? I must know the numbers.
Easy. Just add them. Spicy plus spicy, sweet plus sweet.
[6, 5] Pani Puri
+ [1, 8] Kesari
= [7, 13] the mix
Key idea · Adding and scaling
Adding vectors: add the numbers in the same position. On the drawing, you put the second arrow at the tip of the first. Scaling a vector: multiply every number by the same amount. 2 × [1, 8] = [2, 16] is a double portion. The arrow gets longer but points the same way (just like Super Pongal).
Try it · The mixing counterPick two dishes and the portion size
To do
Bolt's challenge: a customer wants a mix that tastes exactly [7, 9]. Find it.
Adding meanings
Now the magic trick that made AI researchers gasp. If words are vectors, can you add and subtract meanings? Let's give four words two made-up scores: how royal and how male.
Pause and predict
king − man + woman = ?
Aha!Start at king. Subtract man (removes the "male" part). Add woman. You land right on queen. The arrow from man to king means "make it royal", and it works on any word. Try it below.
Try it · Word arithmeticBuild any A − B + C
You just learned
Adding vectors: add the numbers in the same position. [6, 5] + [1, 8] = [7, 13].
Scaling: multiply every number by the same amount. Same direction, longer or shorter arrow.
Meanings can be added too: king − man + woman ≈ queen.
☕
Good place for a break. You've finished the first three chapters, nearly a third of the lesson. Your progress is saved in this browser, so you can close the page and come back any time. The circles at the top show where you stopped.
Chapter 4
Will Priya like it?
Now the big question. Which dish should Bolt give Priya? Ravi knows her well: she loves spice and doesn't like sweet things. He writes her taste as a vector too.
Multiply matching numbers, then add the results. This is called the dot product, written a · b (that little dot comes from Gibbs, the winner of the quaternion war). It gives one number that says how much two vectors agree. And it works the same with 2 numbers or 768: multiply 768 pairs, add them up.
Pause and predict
Grandma Lakshmi dislikes spice and sugar: her taste is [−1, −1]. Which dish will Bolt pick for her?
Aha!Pongal. All scores are negative, but Pongal [2, 1] is the least bad: it's the mildest dish, so the shortest arrow. Bolt picks the highest score, even if it's below zero. Real recommendation systems do exactly this. Check it below.
Try it · Bolt's recommendation enginePick a customer or set a taste
The blue arrow is the taste. Dishes on the arrow's side of the dashed line get a positive score. Dishes exactly on the dashed line score 0.
To do
Bolt's challenge: set Grandma Lakshmi's taste so Bolt's top pick is Pongal.
Bolt notices something. The dot product gets bigger just because an arrow is longer. A plate with double everything scores double. To compare only the direction (the "kind" of taste, remember Super Pongal?), we divide by both lengths. This is called cosine similarity, and it always lands between −1 and 1.
cosine similarity = (a · b) / (|a| × |b|)
1 → same direction (very similar)
0 → at right angles (unrelated)
−1 → opposite (completely different)
🧮Quick refresherAngles: 0°, 90° and 180°Show me
An angle measures how much two arrows turn away from each other, in degrees (°).
0°: they point exactly the same way.
90°: a perfect corner, like the corner of a page. Called a right angle.
180°: they point in exactly opposite directions.
The bars |a| just mean "the length of arrow a".
Try it · The agreement meterDrag either arrow tip
To do
Bolt's challenge: make two arrows that are completely unrelated: dot product exactly 0.
A strange fact about high dimensions
Pause and predict
Pick two arrows completely at random. In 2D, any angle between them is equally likely. What happens in 768 dimensions?
Aha!Almost always very close to 90°: unrelated. There are so many directions to choose from that two random arrows almost never line up. This is great news for AI: unrelated words naturally score near 0. When two embeddings do point the same way, it's no accident; they really are related. See it happen below.
Try it · 2,000 random arrow pairsPick the number of dimensions
In real AISee how it's used
This one small calculation is everywhere:
Search: your question becomes an arrow, every web page becomes an arrow, and the pages pointing most nearly the same way come first.
Chatbots that read your documents: before answering, the chatbot finds the paragraphs most similar to your question this way, then reads them. (People call this RAG, short for "retrieval-augmented generation": look it up, then answer.)
Recommendations: Netflix and YouTube compare your "taste vector" with every video's vector, just like Bolt.
Inside ChatGPT: each word checks which other words in the sentence matter to it by taking dot products. This step is called attention (from a famous 2017 Google paper). You'll build it in a later lesson.
You just learned
The dot product: multiply matching numbers, then add. Big means the two agree; negative means they disagree.
Search, recommendations and chatbots all compare vectors this way.
Chapter 5
A matrix is a machine that moves every dish
Ravi's friend Imran visits from Hyderabad. He eats a Chilli Parotta and laughs: "Anna, this is not spicy at all!" For Imran, every dish feels half as spicy as Ravi's score.
I can make a converter. Ravi's numbers go in, Imran's numbers come out.
Imran's spicy = 0.5 × spicy + 0 × sweet
Imran's sweet = 0 × spicy + 1 × sweet
Write only the numbers: [ 0.5 0 ]
[ 0 1 ] ← this box is a matrix
🧮Quick refresherDecimals like 0.5Show me
A decimal is a number with a dot in it, for parts of a whole. 0.5 is a half, 0.25 is a quarter, 1.5 is one and a half.
So 0.5 × 9 = 4.5: half of 9. Multiplying by 1 changes nothing, and multiplying by 0 gives 0.
Key idea · Matrix
A matrix is a grid of numbers that turns one vector into another. Each row is a recipe for one output number: multiply, then add (a dot product!). And one matrix moves every point at the same time. It can stretch, squash, flip or turn the whole map.
[ a b ] [ x ] [ a·x + b·y ]
[ c d ] × [ y ] = [ c·x + d·y ]
Pause and predict
Here's a matrix: [ 0 −1 ] on top, [ 1 0 ] below. Without calculating anything, where does the step "1 to the right", [1, 0], land?
Aha![0, 1]. Just read the matrix's first column (top to bottom: 0, 1). The first column is always where "1 to the right" lands, and the second column is where "1 up" lands. Every other point simply follows along. So you can read a matrix like a map: this one turns everything a quarter turn. In the lab below you can grab the columns and drag them.
Try it · The converter machineDrag the red and green tips, or type numbers
Faint squares are the old map. Orange lines are the new map. The red tip is where "1 right" lands (column 1) and the green tip is where "1 up" lands (column 2).
To do
Bolt's challenge: find a matrix that squashes every dish onto one straight line. (Not all zeros, that's cheating!) Hint: put the red and green tips on the same line.
New wordDeterminant: how much the machine grows or shrinks every little square. 2 means each square doubles in area. 0 means the squares got flattened into lines, so information was lost and can't be recovered.
The same machine in 3D
With three scores, the matrix is a 3 × 3 grid, and it moves a whole room instead of a page. Its three columns are where "1 along spicy", "1 along sweet" and "1 along sour" land. Watch what it does to a cube.
Try it · A machine for rooms 3DDrag to turn · press a preset
To do
Bolt's challenge: flatten the cube into a flat sheet (volume change 0). Try the presets, or type your own.
In real AI (and games)See how it's used
Every 3D video game moves millions of points with 3 × 3 and 4 × 4 matrices like these, 60 times a second, to turn and move the camera. Remember that; it explains why AI runs on gaming hardware (next chapter).
You just learned
A matrix is a grid of numbers that turns one vector into another.
Each row works like a dot product. One matrix moves every point at once.
Its first column is where "1 right" lands; the second is where "1 up" lands.
Chapter 6
Bolt's brain is a matrix
Ravi asks Bolt to sort dishes into two groups: Breakfast or Dessert. Each dish has 3 scores: spicy, sweet, sour.
Bolt builds a small grid of numbers. 3 numbers go in, 2 scores come out. Each line in the picture below is one number from the grid.
Wait a second…
3 numbers in, multiply-and-add, 2 numbers out. That's the converter machine from Chapter 5! A matrix again, just 2 rows × 3 columns. Keep that thought.
New wordWeight: one number in the grid. It says how much one input matters for one output. Thick line = big weight. Green = positive (pushes the score up), red = negative (pushes it down).
Try it · Train Bolt by handChange the weights
Weights (columns: spicy, sweet, sour)
row 1 → Breakfast · row 2 → Dessert
To do
Bolt's challenge: right now Bolt is wrong. Change the weights so Masala Dosa scores higher for Breakfast and Kesari scores higher for Dessert.
Key idea · A neural network layer
What you just did is one layer of a neural network: output = weights × input. A neural network is many of these layers stacked, each one's output feeding the next (with a small twist in between that you'll learn later). You changed the weights by hand. Training means the computer changes them by itself, a tiny bit at a time, checking how wrong it is after each change, until the answers come out right. An AI model is simply the whole stack with its trained weights.
In real AISee how it's used
GPT-3 is built from many layers like Bolt's, stacked dozens deep. The grids are huge (thousands of rows and columns) and together hold 175 billion weights. Matrix × vector, over and over, that's most of what happens when ChatGPT writes a word.
You just learned
A weight is one number in the grid: how much one input matters for one output.
One neural network layer is output = weights × input: a matrix again.
Training means the computer adjusts the weights itself until the answers come out right.
☕
Good place for a break. Six chapters done. You've already seen what's inside a neural network. Your progress is saved in this browser, so you can close the page and come back any time. The circles at the top show where you stopped.
Chapter 7
Shadows
It's 12 o'clock and the sun is right above the stall. Bolt sees a pole's shadow on the ground and gets an idea.
Priya only cares about one direction: her taste arrow. So for each dish, I only need its shadow on her arrow. The rest she ignores.
Key idea · Projection
The projection of arrow a onto arrow b is a's shadow on b's line. It is "how much of a goes in b's direction". What's left over (the dashed line) always meets b at a perfect right angle. You already did this once: in Chapter 1, looking down from the ceiling threw away the sour number. That was a projection too.
shadow of a on b = (a · b) / (b · b) × b
Pause and predict
Dish a = [6, 5]. Direction b = [1, 0] (pointing right, length exactly 1). How long is a's shadow on b?
Aha!6, and a · b = 6×1 + 5×0 = 6 too. That's no coincidence. When b has length 1, the dot product is the shadow's length. So in Chapter 4, when Bolt scored dishes for Priya, he was secretly measuring shadows on her taste arrow all along! Press "Make b = [1, 0]" in the lab to see it.
Try it · Noon shadowsDrag the dish (red) or the direction (blue)
To do
Bolt's challenge: make the shadow exactly as long as the dish arrow. Nothing left over.
In real AISee how it's used
Drawing the best line through dots (say, predicting a house's price from its size) is a projection. This is called linear regression, and it's often the very first thing people learn in machine learning.
Squeezing data (Pearson's PCA): keep the directions where the data casts the longest shadows and drop the rest. That's how you can draw 768-number embeddings on a flat screen: project them down to 2 or 3 numbers.
You just learned
A projection is one arrow's shadow on another: how much of it goes in that direction.
When the direction has length 1, the dot product is the shadow's length.
Best-fit lines and squeezing data down (PCA) both use shadows.
Chapter 8
Two clues that say the same thing
Ravi's nephew sends an app that gives Bolt two numbers for every dish: chilli in grams and chilli in spoons. Bolt is happy. Two numbers! More information!
Bolt, one spoon is always 5 grams. Those two numbers tell you the same thing twice.
Let's see why that matters. Bolt can mix two "ingredient arrows", any amount of each, and wants to reach a target plate.
Try it · Reach the plateMove the sliders to land on the target
To do
Bolt's challenge: win Level 1, and figure out what's going on in Level 2.
Key idea · Independence and rank
In Level 2, v = 2 × u. They point the same way, so mixing them only ever moves you along one line. We say they are linearly dependent: one is just a stretched copy of the other. Vectors that each add a truly new direction are independent.
The rank of a matrix is how many truly different directions its columns give you. The grams-and-spoons table has 2 columns but rank 1.
★Bonus · optionalThe same trap in 3DOpen bonus
Three arrows in a room should let you reach anywhere in the room. Unless one of them is secretly made from the other two. Below, each dot is a random mix of the three arrows v1, v2 and v3.
Pause and predict
v1 = [1, 0, 0] and v2 = [0, 1, 0]. Now v3 = 2·v1 + v2 = [2, 1, 0]. What shape will all the mixes make?
Aha!A flat sheet: three arrows, but only two real directions (rank 2). No mix can ever lift off the floor to reach [0, 0, 2]. Turn the view in the lab until you see the sheet edge-on: it's paper-thin.
Try it · Ball or sheet? 3DDrag to turn · change v3
To do
Bolt's challenge: start from the ball (press "v3 = [0, 0, 1]"), then change v3 yourself until the cloud collapses into a flat sheet.
New wordFeature: each score you give a computer about a thing. For dishes: spicy, sweet, chilli in grams. For a house: size, rooms, age.
In real AISee how it's used
If two features say the same thing (grams and spoons), the model can't tell which one matters. Its weights become shaky: tiny changes in the data cause wild changes in the answer.
Fine-tuning means taking an AI model that's already trained and teaching it a new skill, like answering in Tamil or in your company's style. Changing every weight is very expensive. In 2021, researchers at Microsoft noticed that the change usually needs only a few truly different directions: it has low rank. Their method, LoRA, stores a change to a 4096 × 4096 grid (16.8 million numbers) as two thin grids of rank 16 (131 thousand numbers). That's over 100 times fewer to train.
You just learned
Two arrows that point the same way only cover one line. They are dependent.
Arrows that each add a new direction are independent.
Rank counts how many truly different directions you really have.
☕
Good place for a break. Only two chapters and the quiz to go. Chapter 9 is a bonus, so you can skip straight to Chapter 10. Your progress is saved in this browser, so you can close the page and come back any time. The circles at the top show where you stopped.
Chapter 9 · Bonus
Cleaning up directions
This whole chapter is optional. It's a neat recipe, but nothing later needs it. Skip to Chapter 10.
★Bonus chapterThe Gram–Schmidt recipeOpen bonus
Bolt has two independent arrows, but they lean on each other. They partly point the same way, so they overlap. Bolt wants two clean directions: each one exactly 1 long, and at a right angle to each other. This tidy pair is called an orthonormal basis ("ortho" = right angle, "normal" = length 1, "basis" = a set of directions you build everything else from). The recipe to make it is the Gram–Schmidt process. It only uses things you already know: shadows and shrinking.
Try it · Step through Gram–SchmidtStep 0 of 4
To do
Bolt's challenge: step all the way to the end.
In real AISee how it's used
Computers use this recipe inside the math tools that solve big systems of equations and fit models, because clean right-angle directions stop rounding errors from piling up.
Chapter 10
Now teach it to Python
Python is a language for giving a computer instructions, one line at a time. It's the most popular language for AI. Everything Bolt learned fits in a few lines of it.
You don't need to install anything. Press ▶ Run under a box and the code runs right here in your browser. You can even change the numbers in the box and run it again. (The first run takes a few seconds while Python loads.)
🐍First time seeing code?How to read these boxesShow me
priya = [2, -1] gives a name to a list. After this line, priya means [2, -1]. (In code, the minus sign is a plain -.)
def dot(a, b): teaches the computer a new word, dot. The indented lines under it are the recipe. return hands back the answer.
print(...) shows something on the screen. Text inside quotes, like "Cosine:", is printed exactly as written.
* means multiply, / means divide, and ** 0.5 means square root.
A line starting with # is a note for humans. The computer ignores it.
round(x, 2) rounds x to 2 digits after the dot.
Vectors, dot product, length, cosine
def dot(a, b):
return sum(x * y for x, y in zip(a, b))
def length(a):
return dot(a, a) ** 0.5
def cosine(a, b):
return dot(a, b) / (length(a) * length(b))
priya = [2, -1]
masala_dosa = [6, 1]
kesari = [1, 8]
print("Priya x Masala Dosa:", dot(priya, masala_dosa))
print("Priya x Kesari: ", dot(priya, kesari))
print("Length of dosa:", round(length(masala_dosa), 2))
print("Cosine:", round(cosine(priya, masala_dosa), 3))
You should seePriya x Masala Dosa: 11
Priya x Kesari: -6
Length of dosa: 6.08
Cosine: 0.809
Same code, 768 dimensions
import random
# two random "words" with 768 numbers each
a = [random.gauss(0, 1) for _ in range(768)]
b = [random.gauss(0, 1) for _ in range(768)]
print("Cosine of two random 768-number arrows:", round(cosine(a, b), 3))
You should seeCosine of two random 768-number arrows: 0.021 (yours will differ, but it'll be close to 0)
A matrix moving a vector, and a neural network layer
def matvec(M, v):
# each row of M does a dot product with v
return [dot(row, v) for row in M]
imran = [[0.5, 0],
[0, 1]]
print("Chilli Parotta for Imran:", matvec(imran, [9, 2]))
# a neural network layer: 3 inputs -> 2 outputs
W = [[ 0.3, -0.2, 0.1], # breakfast
[-0.1, 0.6, -0.1]] # dessert
kesari = [1, 8, 0] # spicy, sweet, sour
print("Kesari scores:", [round(s, 2) for s in matvec(W, kesari)])
You should seeChilli Parotta for Imran: [4.5, 2]
Kesari scores: [-1.3, 4.7]
Shadows (projection)
def project(a, b):
k = dot(a, b) / dot(b, b)
return [k * x for x in b]
shadow = project([3, 4], [1, 0])
print("Shadow of [3, 4] on [1, 0]:", shadow)
You should seeShadow of [3, 4] on [1, 0]: [3.0, 0.0]
💻Later, when you're curiousRunning Python on your own computerShow me
Download Python from python.org/downloads and install it. On Windows, tick "Add Python to PATH" on the first screen.
Open a plain text editor (Notepad on Windows, TextEdit on Mac set to plain text) and paste a snippet. Save it as bolt.py on your Desktop.
Open the Terminal (Mac) or Command Prompt (Windows). Type cd Desktop and press Enter, then python3 bolt.py (on Windows: python bolt.py) and press Enter.
The snippets after the first one use dot and cosine, so paste the first snippet at the top of the same file.
In real AISee how it's used
Engineers use a library called NumPy for this: np.dot(a, b), W @ x, np.linalg.norm(a). It does the same math, just much faster. Now you know what's happening inside it. The full code for this lesson (Python and Julia) is in the course repository.
To do
Bolt's challenge: press ▶ Run on at least one box. Bonus: change a number and run it again.
You just learned
Python turns every idea from this lesson into a few short lines.
dot, length, cosine, matvec and project are all you need.
Real AI code uses NumPy to do the same math much faster.
Final check
Bolt's exam
Ten questions. Get 7 or more right to finish the lesson.
0 / 10
The end (for now)
Bolt's first customer
Ravi anna, the usual?
Your taste is [2, −1]. Chilli Parotta scores 16, the highest. Chilli Parotta for you, Priya.
…How did the robot know?!
Bolt still can't taste. But now it can understand. That's the whole trick of AI: turn things into vectors, and use dot products and matrices to compare them and move them around. In 2 dimensions, 3, or 12,288.
Cheat sheet
Idea
In one line
Where AI uses it
Vector
A list of numbers; also an arrow
Everything a model sees is a vector
Dimension
One kind of score; one direction
Models use hundreds to thousands
Embedding
A learned vector that stands for a thing
Words, images, users
Add / scale
Add same positions; multiply every number
Combining meanings (king − man + woman)
Dot product
Multiply matching numbers, add up
Search, recommendations, attention
Cosine similarity
Compares direction only, −1 to 1
Finding similar documents
Matrix
A grid that moves every vector at once; columns = where the basic steps land
Every neural network layer
Projection
The shadow of one arrow on another
Best-fit lines, PCA
Rank
How many truly different directions
Spotting duplicate features, LoRA
Gram–Schmidt
Turn leaning arrows into clean right-angle ones
Stable equation solving
Next lessonVectors and matrix operations, in more depth: matrix × matrix, transpose, and why the order you multiply in matters.