Text, Sound and Images
Data Representation · 3 question types
Character Sets
Every computer stores text as patterns of binary digits. A character set is the agreed table that maps each character (a letter, a digit, a punctuation mark, an emoji) to a unique binary code. Without a shared character set, a binary value like 01000001 might be interpreted as the letter A on one machine and as a different symbol on another.
A few key ideas about character sets:
- Each character has a unique binary code. No two characters share the same code.
- The codes follow a logical sequence. Consecutive letters (A, B, C) and consecutive digits (0, 1, 2) get consecutive numbers, so the difference between any two adjacent characters is always 1.
- The number of bits decides how many characters can be represented. With n bits there are 2ⁿ possible codes.
The two character sets you must know for the exam are and .
Describing how text is stored using a character set
Question: Describe how text typed at a keyboard is converted to binary, explain how a word is represented using a character set, or name a character set other than ASCII (1–3 marks).
Asked in 4 of the 17 papers. The marks are one per idea: a character set is used, ASCII or Unicode is an example of one, and each character has its own unique binary code. For a whole word, add that the code for each letter is stored in turn, in the order the letters appear. The one-mark versions want only a name: the character set other than ASCII is Unicode, and Unicode is itself an example of a character set.
The point credited in both describe versions is the unique code for each character, so make sure that sentence is in your answer rather than just "the text is turned into binary", which repeats the question.