Learn
1.2.1 | REPRESENTING TEXT
01 | FROM A MESSAGE TO BITS
When you type “Hi”, the computer does not store two tiny pictures of letters as the text itself. It represents the characters using agreed numeric codes, which are stored and processed in binary.
A character set defines a collection of characters and their codes. A character can be a letter, digit, punctuation mark, space or another symbol. Some codes describe control functions rather than printable symbols.
When text is displayed, software interprets the codes and uses a font to draw the characters. The stored character and its visual appearance are related but different things.
02 | A SHARED CODEBOOK
Imagine a class inventing a code where A means 1 and B means 2. A message is useful only if the receiver uses the same rules. Standard character systems let different computers interpret text consistently.
| Character | ASCII denary code | Binary, displayed in a byte |
|---|---|---|
| A | 65 | 01000001 |
| B | 66 | 01000010 |
| a | 97 | 01100001 |
| 0 | 48 | 00110000 |
| Space | 32 | 00100000 |
Uppercase A and lowercase a have different codes because they are different characters. A space also has a code even though it leaves no visible ink on the screen.
You do not need to memorise the entire table. Learn how a mapping works and use a supplied table when answering a question.
03 | ASCII: A SMALL STANDARD CHARACTER SET
ASCII stands for American Standard Code for Information Interchange. Standard ASCII uses seven bits, giving 2⁷ = 128 possible codes, numbered 0 to 127.
It includes English letters, digits, punctuation and control codes. It does not cover the world’s writing systems or emoji.
ASCII values are often displayed or stored in an eight-bit byte with a leading 0. That does not change standard ASCII’s seven-bit code range.
Some older eight-bit systems are called extended ASCII. There is not one universal extended-ASCII table; the interpretation of codes above 127 depends on the particular encoding.
04 | ENCODE AND DECODE A SHORT WORD
| Character | Denary code | Binary byte |
|---|---|---|
| C | 67 | 01000011 |
| A | 65 | 01000001 |
| T | 84 | 01010100 |
Text: CAT Codes: 67 65 84 Binary: 01000011 01000001 01010100
Encoding converts characters into a representation for storage or transmission. Decoding interprets that representation back into characters.
To decode, split this example into bytes, convert each byte to a number and find the character associated with that code. The spaces between groups are just for readability here; they are not extra space characters in CAT.
05 | A TEXT DIGIT IS NOT THE NUMERIC VALUE
The character “5” has ASCII code 53. Its displayed binary byte is 00110101. The unsigned integer value 5 can instead be written as 00000101.
Text character "5": 00110101 (ASCII code 53) Unsigned integer 5: 00000101 (numeric value 5)
The application needs to know what kind of data the bits represent. A string such as “123” contains three characters, not automatically a single binary integer.
Likewise, changing a font does not necessarily change the character codes. A and A in two different fonts can still represent the same character.
06 | UNICODE: MANY LANGUAGES AND SYMBOLS
A global messaging system needs more than English letters. Unicode supports a much greater range of characters and symbols than ASCII, including many writing systems and emoji.
| Character | Unicode code point | What it represents |
|---|---|---|
| A | U+0041 | Latin capital letter A |
| é | U+00E9 | Latin small letter e with acute |
| Ω | U+03A9 | Greek capital letter omega |
| 中 | U+4E2D | A CJK ideograph |
| 😀 | U+1F600 | Grinning face emoji |
A code point is a number assigned in Unicode. U+ is a notation prefix and the following digits are hexadecimal. A code-point label is not itself the sequence of bytes stored in a file.
Unicode preserves the original ASCII character assignments for its first 128 code points. It extends the repertoire rather than giving A a different basic number.
07 | UNICODE AND ITS ENCODINGS
Unicode defines characters and code points. An encoding such as UTF-8 specifies how to represent those code points using bytes.
Do not say that every Unicode character always uses 16 bits. UTF-8 uses one to four bytes for a code point; other Unicode encodings use different rules.
| Character | UTF-8 bytes in hex | Bytes used |
|---|---|---|
| A | 41 | 1 |
| é | C3 A9 | 2 |
| 中 | E4 B8 AD | 3 |
| 😀 | F0 9F 98 80 | 4 |
The ASCII characters each use one byte in UTF-8. Other symbols may use more. A visible symbol can sometimes be made from more than one code point, so counting what appears on screen is not always the same as counting code points or bytes.
For this syllabus, the central comparison is that Unicode can represent a greater range of characters than ASCII. The encoding distinction helps keep that explanation accurate.
08 | WHEN THE READER USES THE WRONG RULES
If a file is encoded one way but decoded using incompatible rules, text can appear as unexpected characters. The binary data needs an agreed interpretation.
For example, UTF-8 encodes é with bytes C3 A9. A reader treating those bytes as Windows-1252 characters can show “é” instead.
A missing font glyph is a different problem: a correctly decoded character may appear as a box if the font or renderer cannot display it. Unicode support does not guarantee every font draws every symbol.
09 | TRY IT: INSPECT YOUR TEXT
Enter a short message. The explorer shows code points and UTF-8 bytes. Try A, a, a space, é and an emoji. Use fictional text; nothing is sent or saved.