We’re moving to a new home! Our website is currently in test mode while we update and transfer our content. Some pages and resources may be temporarily unavailable. We’ll be back with all resources shortly. Thank you for your patience.

CIE IGCSE | 1.3.4 Lossy and Lossless Compression

Lesson objective

Explain lossy and lossless compression, describe how each reduces data size and choose an appropriate method for images, sound, text and other files.

Learn

1.3.4 | LOSSY AND LOSSLESS COMPRESSION

01 | SAME PURPOSE, TWO APPROACHES

A large photograph can be slow to send. A folder of documents can take up valuable storage. Compression represents data using fewer bits, reducing the space needed and, with other factors fixed, the time needed to transfer it.

There are two main approaches. Lossless preserves the information needed to reconstruct the original data exactly. Lossy discards some information to achieve a smaller representation.

Neither approach is automatically best. Choose according to whether exact recovery matters and what quality is acceptable.

02 | LOSSLESS: KEEP ALL THE INFORMATION

Lossless compression finds a more efficient way to represent data. A decoder can reconstruct the original data exactly from the compressed representation.

Imagine sending a row containing ten identical symbols. Instead of writing the same symbol ten times, an agreed format could record the symbol and its repeat count. The information is preserved because the receiver knows how to expand it.

This matters for program code, written documents and database records. Losing one character can change an instruction, a name or a value. ZIP archives and PNG images are familiar examples of lossless compression.

03 | RUN-LENGTH ENCODING: COUNT THE REPEATS

Run-length encoding (RLE) represents consecutive repeated values using a value and a count. In this teaching notation, pairs are written as (symbol, count).

Original: AAAAABBBCC
Encoded:  (A,5) (B,3) (C,2)
Decoded:  AAAAABBBCC

The decoder writes five As, three Bs and two Cs, recovering the original sequence. RLE works well with long runs, such as a simple picture containing large areas of one colour.

The notation is illustrative. Real formats define how symbols and counts are stored. Written punctuation is not a universal RLE file format.

04 | LOSSLESS DOES NOT ALWAYS MEAN SMALLER

Consider a sequence with no repeated neighbours.

Original: ABCDEF
Encoded:  (A,1) (B,1) (C,1) (D,1) (E,1) (F,1)

Now the counts add information rather than saving much repetition. Under a simple model with one byte per symbol and one byte per count, six source bytes become twelve encoded bytes.

By the same model, AAAAABBBCC uses ten source bytes and six encoded bytes. File headers are ignored in both calculations. The effectiveness depends on the data and the method.

Other lossless algorithms use techniques such as dictionaries or efficient codes. RLE is one example, not the method used by every lossless format.

05 | TRY IT: ENCODE AND RECONSTRUCT

Enter up to 200 uppercase letters or digits. The tool groups neighbouring repeats, then reconstructs the original. Compare the sizes under the one-byte symbol plus one-byte count model.

This is a simplified model for learning, not a downloadable compressed file.

06 | LOSSY: DISCARD SELECTED INFORMATION

Lossy compression removes some data. For images, a method may reduce detail or fine colour differences. For sound, it may remove information expected to be less noticeable to a listener.

After decompression, the result is an approximation of the original. The discarded information cannot be recovered from that compressed file alone.

JPEG commonly uses lossy image compression. MP3 uses lossy sound compression. They can be useful when a smaller file matters more than perfect reconstruction.

The amount of visible or audible change depends on the method, settings and source. A lossy file can still look or sound good.

07 | QUALITY AND SIZE: THE TRADE-OFF

More aggressive lossy settings often produce a smaller file at the cost of more distortion. In an image, you may notice blocky areas, blurred detail or marks around sharp edges. In sound, you may notice altered detail or unwanted artefacts.

Do not assume that increasing a quality setting later restores lost detail. Keep an original or lossless master when future editing matters. Repeatedly decoding and re-encoding with lossy methods can introduce further damage.

Simply copying a lossy file does not add compression damage. The risk comes from another lossy encoding step.

08 | A SIMPLIFIED LOSSY MODEL

The samples below are illustrative whole-number values. Round each to the nearest multiple of ten and see what is lost.

Different originals can become the same stored value
Original valueStored rounded value
1210
1410
2730
3330
4850

Both 12 and 14 become 10. A decoder receiving 10 cannot know which original was used. This illustrates information loss, rather than a complete sound compression algorithm.

Rounding alone does not guarantee a smaller file. A real compression scheme must also encode the resulting data more efficiently.

09 | COMPARE THE METHODS

Choose based on purpose, not just the filename
FeatureLosslessLossy
Original recovered exactly?Yes, with correct decodingNo, discarded information is unavailable
How size is reducedMore efficient representation without information lossDiscard information and encode the result efficiently
Quality after decodingOriginal data preservedMay differ from original
Typical useCode, records, text, exact image dataPhotos, audio and media where some loss is acceptable
Size outcomeDepends on data and algorithmOften substantial savings for suitable media; depends on settings

Compression is different from encryption. Compression reduces representation size; encryption protects confidentiality. A ZIP file is not automatically encrypted, and compressed data can still contain sensitive information.

10 | CHOOSE AND JUSTIFY

A program submitted for assessment: use lossless compression so every character can be restored.

A web photograph: lossy compression may provide acceptable quality with a smaller download. Keep the original separately.

A diagram with labels and sharp edges: lossless compression can preserve exact details without lossy artefacts.

An audio master for editing: use lossless or uncompressed storage to preserve source detail. A smaller lossy copy may be suitable for distribution.

In an exam, name the method, connect it to the purpose and explain the consequence. “Lossless because it is better” is too vague.

MATCH THE IMAGE TERM

Terminology

Terminology

Compression

Representing data using fewer bits.

Lossless compression

Compression preserving the information required to reconstruct the original exactly.

Lossy compression

Compression that discards information, preventing exact reconstruction from the compressed file alone.

Decompression

Decoding a compressed representation into usable data.

Run-length encoding

Representing consecutive repeated values using values and repeat counts.

Run

A sequence of adjacent identical values.

Repeat count

The number of times a value occurs consecutively.

Artefact

An unwanted distortion introduced by processing such as lossy compression.

Original / master

A preserved source used to create edited or compressed copies.

Quality trade-off

Balancing acceptable detail against storage and transfer requirements.

File overhead

Additional information, such as headers, that contributes to stored size.

Exact reconstruction

Recovery of the same original data values without information loss.

Questions

Questions

CHECK YOUR UNDERSTANDING

Select all correct choices. Each exact set earns one point.

1. Which statements describe lossless compression?
2. Which statements describe lossy compression?
3. Which notation correctly encodes AAABB using our RLE convention?
4. What sequence does (C,4) (D,2) reconstruct?
5. Which data particularly suits RLE?
6. Which are sensible choices?
7. Which actions can introduce further lossy damage?
8. Which statements are correct?

EXPLAIN AND COMPARE

Write your answers before revealing each sample.

1. Explain the main difference between lossy and lossless compression.

2. Encode WWWWWWWBBGGGG using (symbol, count) pairs.

3. Calculate the size change for that sequence if each symbol and each count uses one byte, ignoring overhead.

4. Explain why RLE may increase the size of ABCDEF under the same model.

5. Recommend compression for a source-code archive and justify your answer.

6. Recommend compression for a photograph on a mobile website and explain a trade-off.

DECISION CHALLENGE | DEFEND YOUR CHOICE

For each scenario, choose lossy or lossless and give a reason: a medical record, a music preview, a labelled revision diagram and a folder of Python files. Compare your choices with a partner.

Reveal suggested reasoning

Medical records and Python files require exact recovery, so use lossless. A music preview may use lossy if acceptable sound quality is retained. A labelled diagram often benefits from lossless storage to preserve crisp text and exact detail.

Flashcards

Flashcards

Click to flip. Select the ideas you need to revisit.

0 cards selected for revision.

    Selections are kept while this page is open.

    Workbook

    Workbook

    COMING SOON

    The workbook for 1.3.4 Lossy and Lossless Compression is coming soon. Complete the encoding tasks and written questions in your exercise book.