Number Systems

From Binary to Hexadecimal – Understanding How Computers Represent Numbers

Per Gustav Ousdal

Introduction

Computers ultimately store bit patterns. A value only gains meaning when those bits are interpreted according to some representation.

This course teaches number systems from that perspective: not only how to convert values on paper, but how binary, hexadecimal, signed integers, masks, endianness and numeric formats appear in real software and hardware.

Prerequisites

The course starts from first principles and assumes no prior knowledge of binary, hexadecimal, assembly language, or digital electronics.

You should be comfortable with basic integer arithmetic: addition, subtraction, multiplication, and division. When powers or other notation are needed, they are introduced before later material depends on them.

Programming and hardware examples are used to show where the representations appear in practice. You do not need prior C, Python, or assembly-language experience to begin the course.

Learning outcomes

By the end of the course you should be able to:

First idea: the number is not its notation

The number twelve can be written in several ways:

The value is the same. The representation is different.

This distinction is the foundation for the rest of the course.

Positional systems and bases

In a positional numeral system, a digit’s value depends on both the digit and its position.

Decimal 347 means 3 × 10² + 4 × 10¹ + 7 × 10⁰.

Binary 1011₂ means 1 × 2³ + 0 × 2² + 1 × 2¹ + 1 × 2⁰ = 11.

Base

The base tells us how many different digits are used before another position is needed.

Base Name Digits
2 binary 0–1
8 octal 0–7
10 decimal 0–9
16 hexadecimal 0–9, A–F

Programming languages often use prefixes such as 0b1011 for binary and 0x2A for hexadecimal. Exact syntax varies between languages and assemblers.

PREDICT

What does 10 mean in base 2, base 8 and base 16?

STEP

In any base b, the notation 10 has the value 1 × b¹ + 0 × b⁰.

OBSERVE

Therefore:

EXPLAIN

The digit sequence alone is not enough. You must also know which base is being used.

Binary: base 2

Binary uses only 0 and 1. Each position has a power-of-two weight.

PREDICT

What decimal value does 10110₂ represent?

STEP

From right to left the weights are 1, 2, 4, 8, 16 and so on. Thus 10110₂ = 16 + 4 + 2 = 22₁₀.

Bits, nibbles and bytes

A bit can hold 0 or 1. Four bits are often called a nibble. Eight bits normally form a byte. Wider machine values are described by an explicit bit width.

With n bits there are 2ⁿ possible patterns. Four bits give 16 patterns; eight bits give 256.

Common misconception

A bit pattern does not have one inherent numeric meaning. 11111111₂ can be 255 as unsigned 8-bit, -1 as signed two’s complement, or part of something that is not an integer at all. The representation rule supplies the meaning.

OBSERVE

For the prediction, 10110₂ = 16 + 4 + 2 = 22₁₀.

Binary makes widths and representation explicit: one more bit doubles the number of available bit patterns.

EXPLAIN

Binary is useful because digital hardware can reliably distinguish two states. The representation maps naturally to logic levels, registers and memory, but the bit pattern still needs context before it has a numeric meaning.

Check yourself

Convert 1111₂ and 10000000₂ to decimal. Explain why eight bits provide 256 patterns but the largest unsigned 8-bit value is 255.

Hexadecimal: base 16

Hexadecimal uses 0–9 and A–F. One hexadecimal digit corresponds exactly to four binary bits.

PREDICT

What binary pattern is represented by A5₁₆?

STEP

A = 10, B = 11, through F = 15. Therefore A5₁₆ = 10×16 + 5 = 165₁₀.

Using four-bit groups:

A5₁₆ = 1010 0101₂.

The mapping is exact because four bits have 2⁴ = 16 possible patterns, exactly matching one hexadecimal digit.

Programming languages commonly use the 0xA5 notation. Assemblers and debuggers may also use forms such as $A5 or an h suffix.

The reverse conversion is equally direct: expand every hexadecimal digit to four bits. For example, 3A₁₆ = 0011 1010₂.

Common misconception

The letters A–F are hexadecimal digits representing the values 10–15. They are not text characters merely because letters are used as symbols.

OBSERVE

For the prediction, A₁₆ = 1010₂ and 5₁₆ = 0101₂, so A5₁₆ = 1010 0101₂.

Two hex digits represent one byte. This makes hex especially convenient for memory dumps, addresses, machine code and bit masks.

EXPLAIN

Hex is compact binary notation. Learning the 16 four-bit patterns removes much of the friction from low-level computing.

Check yourself

Convert FF₁₆ to binary and decimal. Explain why one hex digit always maps to exactly four bits.

Octal: base 8

Octal uses digits 0–7. Each octal digit corresponds exactly to three binary bits.

PREDICT

Translate 111 101 010₂ one three-bit group at a time.

STEP

Base-8 positions have weights 8⁰, 8¹, 8² and so on. For example, 572₈ = 378₁₀, and 101 111 010₂ = 572₈.

Octal remains visible in Unix-style permission notation such as 755, where each digit summarizes three permission bits.

OBSERVE

For the prediction, 111₂ = 7₈, 101₂ = 5₈ and 010₂ = 2₈, so 111 101 010₂ = 752₈.

Three bits represent values 0–7, exactly the digit range of octal.

Common misconception

The value 755 in a Unix file mode is normally read as octal when used as a numeric permission value. It does not mean decimal seven hundred and fifty-five.

EXPLAIN

Octal does not change the number, only its representation. Its computing value comes from its direct three-bit grouping.

Check yourself

  1. Convert 111111₂ to octal.
  2. Convert 17₈ to binary.
  3. Why is the digit 8 invalid inside an octal number?

Converting between bases

Conversion changes representation while preserving value.

PREDICT

Are 1010₂, 12₈, 10₁₀ and A₁₆ different numbers?

STEP

To convert to decimal, sum each digit times its positional weight: 2D₁₆ = 2×16 + 13 = 45₁₀.

To convert a decimal integer to another base, repeatedly divide by the target base and read the remainders in reverse.

Binary provides a convenient bridge. Group four bits for hexadecimal and three bits for octal:

0010 1101₂ = 2D₁₆

101 101₂ = 55₈

OBSERVE

No: 1010₂ = 12₈ = 10₁₀ = A₁₆. The four notations represent the same value in different bases.

The value remains unchanged throughout. Only the symbols and positional weights differ.

EXPLAIN

Conversion becomes much faster once common binary, octal and hexadecimal patterns are recognized directly.

Check yourself

Convert 42₁₀ to binary and hex, and FF₁₆ to decimal.

Powers of two

Computing repeatedly uses quantities of the form 2ⁿ. Knowing common powers of two makes addresses, bit masks and capacities easier to read.

Here n tells us how many factors of 2 are multiplied together. For example, 2³ = 2 × 2 × 2 = 8.

PREDICT

How many distinct values can eight bits represent?

STEP

Useful landmarks are:

n 2ⁿ
0 1
1 2
2 4
3 8
4 16
8 256
10 1024
16 65536
20 1048576
32 4294967296

With n bits there are 2ⁿ bit patterns. Eight bits therefore provide 256 patterns, representing 0 through 255 when interpreted as unsigned.

KiB and kB

1 KiB = 1024 bytes, while the SI quantity 1 kB = 1000 bytes. MiB and GiB are corresponding binary prefixes.

Common misconception

2⁸ = 256 means that eight bits provide 256 different patterns. It does not mean that the largest unsigned value is 256. Counting starts at zero, so the largest value is 255.

OBSERVE

Eight bits form 2⁸ = 256 distinct bit patterns and can therefore represent 256 distinct values once an interpretation is chosen.

One hexadecimal digit is four bits, two hexadecimal digits are one byte, and four hexadecimal digits span 16 bits.

EXPLAIN

Powers of two connect bit width, ranges, memory sizes and address spaces.

Check yourself

  1. How many patterns fit in 16 bits?
  2. What is the largest unsigned 8-bit value?
  3. What is the difference between 1 MiB and 1 MB?

Unsigned integers

An unsigned integer uses every bit to represent a non-negative magnitude.

PREDICT

What is the largest value that fits in eight bits?

STEP

With n bits there are 2ⁿ patterns, giving the unsigned range 0 ... 2ⁿ − 1.

Examples:

Width Range
8 bits 0–255
16 bits 0–65,535
24 bits 0–16,777,215
32 bits 0–4,294,967,295

11111111₂ = FF₁₆ = 255₁₀ when interpreted as unsigned eight-bit data.

Overflow and wraparound

Fixed-width arithmetic cannot represent every integer. Under modulo-2ⁿ arithmetic, keeping only eight bits makes 255 + 1 wrap to 0.

Common misconception

Eight bits provide 256 possible patterns, but the unsigned range is 0–255 because zero consumes one of those patterns.

OBSERVE

For an unsigned 8-bit integer, the maximum is 2⁸ − 1 = 255, or 11111111₂ = FF₁₆.

The bits alone do not say “unsigned”. Their interpretation depends on the type and width supplied by context.

EXPLAIN

Width and interpretation matter when reading registers, protocols, file formats and machine code.

Check yourself

Find the range of a 12-bit unsigned integer and compute 250 + 10 with eight-bit wraparound.

Signed integers

Negative integers require an agreed interpretation of bit patterns.

PREDICT

Can 11111111₂ mean both 255 and −1?

STEP

Sign-magnitude reserves a sign bit and therefore has both +0 and −0.

Ones’ complement forms a negative value by inverting every bit and also has two zero encodings.

Two’s complement is the standard modern integer representation. For n bits its range is −2ⁿ⁻¹ ... 2ⁿ⁻¹ − 1. Eight bits therefore cover −128 through 127.

Common misconception

The top bit in two’s complement is not merely a separate minus flag that can be removed from the rest of the number. The entire bit pattern participates in the representation.

This is also why the range is asymmetric: an eight-bit two’s-complement integer has one more negative value than positive values.

OBSERVE

The same byte FF₁₆ is 255 as unsigned eight-bit data and −1 as signed eight-bit two’s complement.

EXPLAIN

Meaning requires context: width, signedness and encoding rule. Raw bits do not carry those labels themselves.

Check yourself

What is the signed 16-bit two’s-complement range? Explain the asymmetry around zero.

Two’s complement in practice

Two’s complement lets binary addition hardware operate naturally on positive and negative integers.

PREDICT

How can 11111011₂ represent −5?

STEP

To encode −5 in eight bits, write +5 as 00000101, invert the bits to 11111010, then add one to obtain 11111011 or FB₁₆.

For an n-bit pattern whose top bit is one, another interpretation method is unsigned-value minus 2ⁿ: 251 - 256 = -5.

Addition

00000101 + 11111011 = 1 00000000. Discarding the carry beyond eight bits leaves zero.

The minimum-value edge case

Negation is invert-plus-one, but signed eight-bit −128 has no representable +128 counterpart. Code must account for this boundary case.

OBSERVE

For the prediction, 11111011₂ is −5 in 8-bit two’s complement: invert 00000101₂ to 11111010₂ and add 1 to obtain 11111011₂.

Carry and signed overflow answer different questions, so CPUs may expose separate flags for them.

EXPLAIN

Two’s complement is naturally understood as modulo-2ⁿ arithmetic plus a signed interpretation of the bit patterns.

Check yourself

Encode −1, −42 and −128 as eight-bit hexadecimal. Interpret 80₁₆ as both unsigned and signed.

Binary arithmetic and CPU flags

Binary addition follows ordinary positional arithmetic with only two digits.

PREDICT

What is 1111₂ + 0001₂ in a four-bit register?

STEP

1+1=10₂, producing a carry. Thus 1111 + 0001 = 1 0000. Keeping four result bits gives zero while the carry records unsigned overflow beyond the width.

Subtraction can use borrow directly or addition of a two’s-complement negative operand.

CPU flags

Many CPUs record properties of the result. Common flags include:

For signed eight-bit arithmetic, 7F₁₆ + 01₁₆ = 80₁₆. The mathematical +128 cannot be represented, so CPUs with an overflow flag report that condition.

Common misconception

Flag names and exact behavior are not identical across CPU architectures. The instruction-set documentation defines which flags an instruction changes and what they mean there.

Carry and signed overflow are also different conditions: an operation can set one without setting the other.

OBSERVE

With only four bits, 1111₂ + 0001₂ = 1 0000₂. The register retains the low four bits 0000₂, while the fifth bit is the carry out.

Carry and signed overflow answer different questions about the same bit-level result.

EXPLAIN

Flags allow later instructions to branch, compare and extend arithmetic without recomputing the operation.

Check yourself

Compute FF₁₆ + 01₁₆ and discuss C, Z and V for eight-bit arithmetic.

Bitwise operations

Bitwise operations process corresponding bits independently.

PREDICT

What is 1010 AND 1100?

STEP

AND requires both bits to be one. OR requires at least one. XOR is one when the bits differ. NOT inverts every bit.

For A=1010₂, B=1100₂: AND gives 1000, OR 1110, XOR 0110, and four-bit NOT A gives 0101.

Logical shifts move bits and fill with zero. Arithmetic right shift commonly preserves the sign bit. Rotates wrap bits around instead of discarding them; rotate-through-carry also includes the carry flag.

OBSERVE

For the prediction, compare corresponding bits: 1010 AND 1100 = 1000.

Width is essential: NOT zero has a different numerical result at 8 and 16 bits.

EXPLAIN

These operations underpin register manipulation, protocols, graphics, compression and many low-level algorithms.

Check yourself

Compute AND, OR and XOR for 3C₁₆ and 0F₁₆.

Bit masks and bit fields

A bit mask selects particular bits while leaving neighboring bits available for other meanings.

PREDICT

How can bit 0 be set in 10100100₂ without changing the other bits?

STEP

With MASK=00000001₂: set using OR, clear using AND with NOT MASK, toggle using XOR, and test using AND.

Fields spanning several bits can be isolated with a mask and shifted down before interpretation.

Suppose bit 7 means ENABLE, bit 3 IRQ and bit 0 READY. Then 10001001₂ carries three independent Boolean flags in one byte.

OBSERVE

Set bit 0 by OR-ing with 00000001₂: 10100100₂ OR 00000001₂ = 10100101₂. The other bits are unchanged.

Hex makes masks compact: 11110000₂ = F0₁₆, while 00001111₂ = 0F₁₆.

EXPLAIN

Bit fields efficiently pack small values, and masks let software manipulate one field without damaging the rest.

Check yourself

Create a mask selecting bits 2 and 5, then show how to toggle bit 5 in 34₁₆.

Memory, addresses and offsets

Memory can be modeled as a sequence of addressable bytes. An address identifies a location; a value may occupy one or several consecutive bytes.

PREDICT

If a 32-bit value starts at address 0x1000, which byte addresses does it occupy?

STEP

In byte-addressable memory the address normally increases by one per byte: 0x1000, 0x1001, 0x1002, 0x1003. A 32-bit value occupies four bytes.

An offset is a displacement relative to a base address. Base 0x2000 plus offset 0x34 gives 0x2034.

With n address bits, 2ⁿ address patterns exist. On a byte-addressed system, 16 address bits can identify 65,536 byte locations.

Alignment

Architectures may prefer or require multi-byte values to begin at particular address boundaries. A 32-bit value may be naturally aligned at an address divisible by four.

OBSERVE

The predicted 32-bit value occupies four bytes: 0x1000, 0x1001, 0x1002 and 0x1003.

Hexadecimal maps cleanly to address bits: a 16-bit address is four hex digits from 0000 through FFFF.

EXPLAIN

Debugger and hex-editor displays require you to distinguish the address of data from the value stored there.

Check yourself

Which addresses are occupied by a 16-byte block beginning at 0x3FF0? Compute base 0x8000 plus offset 0x2A.

Endianness

Endianness describes the byte order used to store a multi-byte value. It does not change the numerical value itself.

PREDICT

How can 0x12345678 occupy four bytes beginning at 0x1000?

STEP

Big-endian stores the most significant byte first: 12 34 56 78 at increasing addresses.

Little-endian stores the least significant byte first: 78 56 34 12.

The bit notation inside each normally written byte is not reversed. Here, endianness is about ordering bytes of a multi-byte value.

The Motorola 68000 family is associated with big-endian storage, while x86 uses little-endian. Protocols and file formats can define byte order independently of the host CPU.

Memory dumps

A dump showing 78 56 34 12 can encode the 32-bit number 0x12345678 when the format specifies little-endian. Raw bytes need width, type and byte-order context.

OBSERVE

For the prediction, 0x12345678 may occupy memory as 12 34 56 78 in big-endian or 78 56 34 12 in little-endian order from the lowest address. Byte order is part of the representation.

A hex dump normally displays bytes in increasing memory-address order, which need not match the written digit order of a multi-byte number.

EXPLAIN

Endianness matters when exchanging data among CPUs, networks, binary files, emulators and debugging tools.

Check yourself

Write 0xA1B2C3D4 as four bytes in both big- and little-endian order.

BCD — Binary-Coded Decimal

BCD encodes decimal digits individually rather than representing the entire number as one binary integer.

PREDICT

How can decimal 42 be stored when each decimal digit receives four bits?

STEP

Packed BCD gives each digit one nibble: 42 -> 0100 0010₂ -> 0x42. Ordinary binary 42 is instead 00101010₂ = 0x2A.

Only nibble values 0–9 are normally valid decimal digits. Unpacked BCD commonly gives each digit a whole byte.

BCD has been useful in calculators, clocks, counters, financial systems and historical computers where decimal digits need direct preservation.

OBSERVE

With four bits per decimal digit, 4 is 0100 and 2 is 0010. Packed BCD for 42 is therefore 0100 0010₂ = 0x42.

A byte 0x42 can mean ordinary integer 66 or packed-BCD decimal 42. The format supplies the meaning.

EXPLAIN

BCD reinforces a central theme: bits need a representation rule before they become a number.

Check yourself

Encode 1987 as packed BCD and state its storage size.

Fixed-point and Q formats

Fixed-point represents fractional values using an implied, fixed binary-point position.

PREDICT

With four integer bits and four fractional bits, what does 00111000₂ represent?

STEP

Interpreted as 0011.1000₂, the value is 3.5. The stored integer is 56 and four fractional bits imply division by 2⁴ = 16.

With f fractional bits, the resolution is 1 / 2^f.

Two’s-complement signed integers can use the same scaling rule for signed fixed-point values.

Tradeoffs

Fixed-point offers predictable resolution and efficient arithmetic on systems without fast floating-point hardware, but software must manage scaling, range and overflow explicitly.

OBSERVE

With four fractional bits, place the binary point after the upper four bits: 0011.1000₂ = 3 + 1/2 = 3.5.

No physical binary-point symbol is stored. Its location belongs to the format definition.

EXPLAIN

Fixed-point is scaled integer arithmetic, directly connecting fractional numbers to integer width and overflow.

Check yourself

What is the resolution with eight fractional bits? Interpret stored integer 384 using that scale.

Floating-point and IEEE 754

Floating-point stores a significand together with an exponent so the binary point can effectively move.

PREDICT

How can the same 32-bit format cover both very large and very small magnitudes?

STEP

IEEE 754 binary32 contains one sign bit, eight exponent bits and 23 explicit fraction bits. Normal values conceptually follow (-1)^sign × significand × 2^exponent; the exponent uses a bias and the leading one of a normal binary significand is implicit.

IEEE 754 also defines signed zeros, infinities, NaNs and subnormal values.

Worked example: 1.0 = 0x3F800000

The binary32 pattern 0x3F800000 is:

0 01111111 00000000000000000000000

The value is therefore (+1) × 1.0₂ × 2⁰ = 1.0.

From a decimal fraction to a binary fraction

A fraction can be converted by repeatedly multiplying its fractional part by 2 and recording the integer part:

For 0.1₁₀, the process does not terminate: the bit pattern repeats. There is therefore no finite binary fraction exactly equal to one tenth.

Why 0.1 + 0.2?

One tenth has no finite binary fractional expansion, just as one third has no finite decimal expansion. It must be rounded to a representable binary floating-point value. Arithmetic on rounded operands can therefore differ slightly from ideal decimal arithmetic. In ordinary binary64 arithmetic, for example, 0.1 + 0.2 is typically represented as approximately 0.30000000000000004, not exactly 0.3.

Precision and range

Exponent bits provide range while significand bits provide precision. Floating-point is a controlled compromise, not arbitrary-precision real arithmetic.

OBSERVE

0x3F800000 demonstrates concretely how the sign, exponent and fraction fields combine to produce 1.0. The 0.1 example also shows why not every decimal fraction can be stored exactly.

A floating-point bit pattern cannot be interpreted as an ordinary integer while preserving its numerical meaning; the field structure defines the representation.

EXPLAIN

IEEE 754 standardizes behavior, but programs still need to account for rounding, special values and comparisons.

Check yourself

Why is exact floating-point equality often more subtle than integer equality?

Other representations

Bit patterns represent much more than ordinary integers. This chapter introduces Gray code, biased/excess notation, character codes and packed data.

PREDICT

Must the bit pattern 01000001 always mean the number 65?

STEP — Gray code

In binary-reflected Gray code, neighboring values differ in exactly one bit. This is useful when multiple bits should not transition simultaneously, such as in some position encoders.

A binary value b can be converted with g = b XOR (b >> 1).

STEP — biased/excess notation

An exponent can be stored as an unsigned field plus an agreed bias. With bias 127, stored value 130 represents exponent (130-127=3).

IEEE 754 binary32 uses this principle for its exponent field, with reserved field values receiving special meanings.

STEP — characters as numbers

ASCII assigns A the code value 65, or 0x41. Unicode defines code points, such as U+0041 for A. An encoding such as UTF-8 then determines the actual stored bytes.

A code point and its byte encoding are not the same concept.

STEP — packed data

One byte can contain several bit fields. A hypothetical status register might use bit 7 for ready, bit 6 for error, bits 5–4 for a mode and bits 3–0 for a counter.

OBSERVE

No. 01000001₂ = 0x41 is unsigned integer 65, but the same pattern can, for example, represent the ASCII character A.

0x41 may be unsigned 65, signed 65, ASCII A, part of an instruction or a field inside a larger structure. Context supplies meaning.

EXPLAIN

Computers store bits. Types, protocols, instruction sets and file formats supply their semantics.

Check yourself

Compute the Gray code for binary 1010. What exponent does stored value 124 represent with excess-127?

Numbers in assembly

Assembly exposes the distinction among value, syntax, bit width and machine meaning. A number may be an immediate constant, address, mask, offset or part of an instruction encoding.

PREDICT

Do $10, #10, 0x10 and 10h always mean the same thing?

STEP — literals and syntax

Assembler dialects use different literal conventions. Common forms include decimal 42, hexadecimal 0x2A, $2A in many classic assemblers, 2Ah in some Intel-style contexts, and forms such as %101010 for binary.

Always interpret syntax in the context of the actual assembler.

Immediate versus address

A numeric operand can be the value itself or identify a location from which a value is read.

On 6502-style syntax, LDA #$2A commonly loads immediate value 0x2A, while LDA $2A accesses memory according to the instruction’s addressing mode.

On Motorola 68000, MOVE.B #$2A,D0 uses an immediate value, while MOVE.B $002A,D0 refers to memory. Size suffixes such as .B, .W and .L make operand width visible in many 68k dialects.

AVR

AVR uses its own register and instruction constraints. For example, LDI R16, 0x2A loads a constant into a permitted register. A source literal still has to fit the operand range and instruction encoding.

x86

x86 has multiple major assembler dialects. Intel and AT&T syntax can express the same machine operation differently, so no single textual convention should be mistaken for universal assembly syntax.

Registers, masks and disassembly

Registers and instruction fields have finite widths. Bit masks such as 0x0F are commonly used to select fields.

A disassembler interprets raw bytes as instructions for a chosen ISA and address. Bytes that are actually data can look like plausible instructions when context is wrong.

OBSERVE

No. Prefixes and suffixes are assembler syntax, and the meanings of `# Numbers in assembly

Assembly exposes the distinction among value, syntax, bit width and machine meaning. A number may be an immediate constant, address, mask, offset or part of an instruction encoding.

PREDICT

Do $10, #10, 0x10 and 10h always mean the same thing?

STEP — literals and syntax

Assembler dialects use different literal conventions. Common forms include decimal 42, hexadecimal 0x2A, $2A in many classic assemblers, 2Ah in some Intel-style contexts, and forms such as %101010 for binary.

Always interpret syntax in the context of the actual assembler.

Immediate versus address

A numeric operand can be the value itself or identify a location from which a value is read.

On 6502-style syntax, LDA #$2A commonly loads immediate value 0x2A, while LDA $2A accesses memory according to the instruction’s addressing mode.

On Motorola 68000, MOVE.B #$2A,D0 uses an immediate value, while MOVE.B $002A,D0 refers to memory. Size suffixes such as .B, .W and .L make operand width visible in many 68k dialects.

AVR

AVR uses its own register and instruction constraints. For example, LDI R16, 0x2A loads a constant into a permitted register. A source literal still has to fit the operand range and instruction encoding.

x86

x86 has multiple major assembler dialects. Intel and AT&T syntax can express the same machine operation differently, so no single textual convention should be mistaken for universal assembly syntax.

Registers, masks and disassembly

Registers and instruction fields have finite widths. Bit masks such as 0x0F are commonly used to select fields.

A disassembler interprets raw bytes as instructions for a chosen ISA and address. Bytes that are actually data can look like plausible instructions when context is wrong.

, #, 0x and h depend on the language and context. The syntax rules of the particular assembler are authoritative.

Assembly source contains human-readable symbols and numeric notation. The CPU receives encoded instruction bit fields.

EXPLAIN

Number representation is fundamental to assembly because operands, addresses, registers, masks, offsets and encodings are all finite bit fields.

Check yourself

Why are #$2A and $2A different in many classic assembler dialects? Why must you know the dialect before interpreting 10h?

Numbers in C and Python

C and Python can express the same binary and hexadecimal values, but their integer models differ substantially.

PREDICT

Will 255 + 1 always behave the same way in C and Python?

STEP — literals

Hexadecimal and decimal literals such as 0x2A and 42 work the same way in both languages. Binary literals such as 0b101010 are supported in Python and, from C23, in C (many compilers, including GCC and Clang, have long accepted them as an extension). In C, language type and literal rules determine the resulting type. Python int is not confined to a fixed 8-, 16-, 32- or 64-bit width.

C: explicit widths

When machine width matters, <stdint.h> provides types such as uint8_t, uint16_t and int32_t when the implementation has suitable exact-width integer types.

Unsigned C arithmetic wraps modulo 2ⁿ for the type width. Do not assume the same rule for signed overflow; signed integer overflow is not defined by C as ordinary two’s-complement wraparound.

Python: arbitrary-size integers

Python integers grow as required within practical resource limits. To model an 8-bit register explicitly, impose a width:

x = (255 + 1) & 0xFF

The result is zero.

Shifts, masks, parsing and formatting

Bitwise operations look familiar in both languages, but type widths, promotions and negative-value behavior require language-specific care.

Python can parse bases with int(text, base) and format values with forms such as f"{42:08b}" and f"{42:02X}".

C offers conversion routines such as strtoul; portable formatting of fixed-width integers can use the macros in <inttypes.h>.

OBSERVE

No. The result depends on type and language rules. Python integers normally grow as needed, while arithmetic involving bounded C integer types follows different rules. Both the type and language semantics matter.

The source literal 0xFF is a numeric value. Storing or interpreting it as an 8-bit signed value, 8-bit unsigned value or wider type is a separate operation.

EXPLAIN

Programming languages add rules on top of bit representations. Reliable low-level code requires understanding both the machine model and the language model.

Check yourself

Why does (255 + 1) & 0xFF produce zero in Python? Why should signed C overflow not be assumed to work the same way?

Numbers in digital electronics

Digital electronics connects bit patterns to physical signals, registers, measurements and protocols. Numbers do not merely describe mathematics: they control pins and encode observations of the physical world.

PREDICT

If an 8-bit GPIO register contains 0x81, which bits are set?

STEP — logical and physical levels

Logical zero and one are abstractions. Physical circuits use voltage ranges interpreted as LOW and HIGH; exact thresholds come from the component datasheet.

Signals can be active-low, meaning the asserted function corresponds to an electrically LOW line. Naming conventions vary.

GPIO registers and masks

For PORT = 10100001₂ = 0xA1, individual bits may control separate functions.

Typical software operations include setting a bit with PORT |= (1 << 3), clearing it with PORT &= ~(1 << 3), and toggling it with XOR.

Real hardware must be checked against its datasheet. Some registers provide dedicated SET/CLEAR aliases, write-one-to-clear flags or other semantics for which ordinary read-modify-write is inappropriate.

Datasheets and register maps

Register documentation commonly specifies addresses or offsets, bit positions, field widths, reset values, access properties and meanings of encoded field values.

A value such as 0x82 is only meaningful once its register and field definitions are known.

ADC and DAC

An idealized N-bit ADC provides 2^N digital codes. A 10-bit converter has 1024 codes, usually numbered 0 through 1023.

A simple idealized unipolar conversion can be approximated by

V ≈ code / (2^N - 1) × Vref

but the exact transfer function, reference, tolerances and endpoint conventions come from the datasheet. The formula above is therefore a learning model, not a universal ADC law.

A DAC maps digital codes in the opposite direction toward an analog output; its width likewise determines the number of available codes.

Logic analyzers

Logic analyzers capture digital samples and may display bits, bytes, hexadecimal values, timestamps or decoded UART/SPI/I²C fields.

The byte 10100101₂ = 0xA5 might be a command, address, flags or payload. Protocol context supplies meaning. Serial analysis may additionally require bit order, clocking, framing and multi-byte byte order.

OBSERVE

For the prediction, 0x81 = 1000 0001₂, so bits 7 and 0 are set.

Hardware registers and protocols demonstrate the course’s central idea: bit pattern plus representation rules produces meaning.

EXPLAIN

Reading a datasheet constantly requires translating among bit positions, masks, hexadecimal values, numerical quantities and physical behavior.

Check yourself

Which bits are set in 0x81? How many codes does a 12-bit ADC have? Why can ordinary read-modify-write be wrong for some status registers? Why must an ADC conversion formula be checked against the datasheet?

Numbers in networks and file formats

Network packets and binary files are sequences of bytes. Understanding them requires knowing which bytes form fields, the byte order of multi-byte values and the meaning assigned by the format.

PREDICT

The bytes 12 34 form a 16-bit integer. Is the value 0x1234 or 0x3412? Write down your hypothesis before continuing.

STEP — network byte order

Many multi-byte fields in Internet protocols use network byte order, which is big-endian. A 16-bit 0x1234 field is therefore transmitted as 12 34.

This does not require the host CPU itself to be big-endian. Protocol representation and native machine representation are separate concerns.

IPv4 and MAC addresses

IPv4 is 32 bits. The familiar 192.168.1.10 corresponds to bytes C0 A8 01 0A.

A common 48-bit MAC address can be displayed as six hexadecimal bytes, for example 02:12:34:56:78:9A. Hexadecimal maps naturally to byte-oriented representations.

Protocol fields

Packets may contain versions, lengths, types, flags, ports, sequence numbers, checksums and payloads. Some are bit fields; others are multi-byte integers. The protocol specification defines width, byte order and semantics.

Magic numbers and binary formats

Binary formats often begin with characteristic signatures. For example, 89 50 4E 47 0D 0A 1A 0A is the eight-byte signature at the start of a PNG file.

A signature can aid identification, but does not by itself prove that the rest of a file is structurally valid.

From hex dump to structure

Consider a fictional record:

01 03 12 34 41 42 43

Its specification says byte 0 is a version, byte 1 a payload length, bytes 2–3 a big-endian ID and the remainder payload.

The record therefore contains version 1, length 3, ID 0x1234, and payload 41 42 43, which can be interpreted as ASCII ABC.

Without the specification these are simply seven bytes.

Offsets and lengths

Binary parsers move through structures using field offsets and lengths. A field beginning at offset 4 with length 2 occupies bytes 4 and 5; the next field begins at offset 6.

An incorrect width or offset shifts subsequent interpretation.

OBSERVE

There is no single answer to the prediction without a byte-order rule: big-endian gives 0x1234, while little-endian gives 0x3412.

The same byte can be part of an IP address, character, flags, integer or opaque payload.

EXPLAIN

A hex dump exposes stored or transmitted bytes. A format or protocol specification tells us how to group and interpret them.

Check yourself

How is 0xBEEF stored as a big-endian 16-bit field? Which bytes encode IPv4 127.0.0.1? Why is a magic number insufficient to validate an entire file?

Numbers in debugging and reverse engineering

When software crashes, a file is undocumented or an old system lacks specifications, raw bytes may appear before their structure is known. Number representation becomes an analysis tool.

PREDICT

You find the bytes 48 65 6C 6C 6F 00. Are they integers, machine code or text?

STEP — observe before assuming

Separate observations from hypotheses. A hex editor or memory dump may show offsets or addresses, hexadecimal bytes, a text column and highlighted regions.

Readable characters are evidence worth investigating, not proof that a field is text.

Recognizing patterns

Useful observations include repeated zeros, recurring values, text-like sequences, possible lengths or offsets, known signatures, regular blocks and bit fields whose bits change predictably.

A hypothesis becomes stronger when it explains multiple independent examples.

Integers, addresses and offsets

Bytes 34 12 may represent big-endian 0x3412 or little-endian 0x1234.

If the field is hypothesized to be a length, test whether that value agrees with the surrounding structure.

Debugger addresses are commonly hexadecimal. Address differences reveal sizes: 0x1040 - 0x1000 = 0x40 = 64.

Disassembly

A disassembler interprets bytes as instructions for a selected CPU and starting point. That does not prove those bytes are executable code.

Data can decode into syntactically valid but meaningless instructions. Ask whether the CPU and start address are known, whether control flow reaches the region, whether the sequence is plausible and whether the bytes might instead be data.

Code and data

Code, tables, strings and constants may live close together. 41 42 43 44 could be ASCII ABCD, four integers, portions of larger values or instruction bytes. Context determines the useful interpretation.

Reconstructing a possible structure

Consider the fictional bytes 03 00 08 00 41 42 43 00.

One hypothesis is type 3, flags 0, a little-endian length of 8 and four data bytes. This is a model, not yet a fact. Comparing additional records can test whether those fields behave consistently.

Change one thing

With systems and test data you control, changing one known value at a time and comparing before and after is powerful.

If enabling a flag changes exactly one bit, that is useful evidence. If changing a counter tracks one field, its representation becomes clearer.

OBSERVE

Several interpretations are possible without context. As ASCII/UTF-8, 48 65 6C 6C 6F 00 gives Hello followed by a zero byte; that is a strong hypothesis, not proof by itself.

Data reverse engineering is often a process of constructing hypotheses that can be weakened or strengthened rather than instantly guessing the correct format.

EXPLAIN

Bases, signedness, byte order, bit fields, addresses and text encodings make raw bytes easier to reason about. Context and repeated observations distinguish plausible interpretation from documented structure.

Check yourself

Why can data produce apparently valid disassembly? What could 00 01 mean under different byte orders? Why are multiple examples more useful than one dump when reconstructing a structure?

Historical and unusual number bases

Binary, octal, decimal and hexadecimal dominate modern computing, but positional notation can use many other bases. Exploring them reinforces that a base is a representation rule rather than a property of the number itself.

PREDICT

What do you think 10 means in base 3? Write down your answer before continuing.

STEP — base 3

Base 3 uses digits 0, 1 and 2. For example, 102₃ = 1 × 9 + 0 × 3 + 2 = 11₁₀.

Ternary systems demonstrate that digital representation is not mathematically restricted to exactly two symbols. Ternary logic has also appeared in computing research and historical machines.

Base 12

Base 12 requires digit values for ten and eleven in addition to 0–9. In this course we use A for ten and B for eleven, so A₁₂ = 10₁₀ and B₁₂ = 11₁₀. Twelve has several divisors — 2, 3, 4 and 6 — so some fractions are compact.

One half in base 12 is 0.6₁₂, because 6/12 = 1/2.

Base 20

Vigesimal systems use base 20 and occur historically in multiple languages and cultures. For representation, the key point is simple: 10₂₀ means twenty rather than ten.

Base 36

Base 36 can use 0–9 and A–Z as its 36 digit values. Thus Z₃₆ = 35₁₀ and 10₃₆ = 36₁₀.

It can provide compact textual forms for non-negative integers, although a format must define its alphabet and case rules.

Base 60

Sexagesimal systems have ancient roots. Base-60 structure remains visible in time and angle measurement: 60 seconds per minute, 60 minutes per hour, 60 arcminutes per degree and 60 arcseconds per arcminute.

These conventions are not all simply modern positional base-60 notation, but they illustrate how a numerical subdivision can persist.

Why divisibility matters

Factors of the base affect which fractions terminate.

In base 10, 1/2 and 1/5 terminate while 1/3 repeats. In base 2, 1/2 terminates but decimal 1/10 does not have a finite binary fractional representation. This connects directly to floating-point behavior.

Base is not storage width

A base and a storage width are separate concepts. A value displayed as base-36 text may be stored internally as an ordinary binary integer. A byte can likewise be displayed in decimal, hexadecimal or another base without changing its value.

Historical machines

Computing history includes representations unlike today’s most familiar binary conventions. Some machines emphasized decimal arithmetic or used unusual word sizes and character encodings.

Historical data must therefore be interpreted according to the documented representation of the relevant machine rather than modern assumptions.

OBSERVE

For the prediction, 10₃ means three: 1 × 3¹ + 0 × 3⁰ = 3.

The notation 10 can represent 2, 3, 8, 10, 12, 16, 20, 36, 60 or another value depending on the base.

EXPLAIN

A positional system defines digits, positional weights and a radix. The number is the abstract value; the notation is its representation.

Check yourself

What is 10₁₂ in decimal? What is Z₃₆? Why can base 12 express some common fractions more compactly than base 10?

Final project: analyze an unknown byte buffer

This project combines the ideas from EduNumbers. You will use representation, byte order, signedness, bit fields, text and structure to explain a synthetic unknown buffer.

The task

0000: 45 4E 55 4D 01 A5 10 00 34 12 FE FF 50 4C 4F 4F
0010: 53 00 00 00 78 56 34 12

You initially know only that the format was created for this exercise, every field starts on a byte boundary, and the buffer is complete.

PREDICT

Before calculating, identify regions that look like text, flags or possible little-endian integers. Record hypotheses first.

STEP — build evidence

The first four bytes decode as ASCII ENUM, making them a strong signature candidate.

Byte 0x01 at offset 0x04 is plausible as a version. Byte 0xA5 at 0x05 is 10100101₂, with bits 7, 5, 2 and 0 set.

Bytes 10 00 become 16 as little-endian 16-bit; as big-endian they would be 4096. Since the complete buffer contains 16 bytes after its eight-byte header, 16 is a plausible length, while 34 12 becomes 0x1234 = 4660.

Bytes FE FF become 65534 unsigned or -2 as signed 16-bit two’s complement. The bits alone do not specify signedness.

Bytes 50 4C 4F 4F 53 00 00 00 resemble an eight-byte null-padded ASCII field containing PLOOS.

Finally, 78 56 34 12 becomes 0x12345678 as little-endian. Multiple consistent fields strengthen the little-endian hypothesis.

A consistent structure

Offset Size Possible field Interpretation
0x00 4 magic ASCII ENUM
0x04 1 version 1
0x05 1 flags 0xA5
0x06 2 length 16, little-endian
0x08 2 id 0x1234
0x0A 2 delta -2 signed
0x0C 8 name PLOOS, null-padded
0x14 4 value 0x12345678

For this synthetic exercise that is the intended structure. In real reverse engineering, more samples or documentation would be required before treating such a model as established fact.

Verify with code

data = bytes.fromhex(
    "45 4E 55 4D 01 A5 10 00 34 12 FE FF "
    "50 4C 4F 4F 53 00 00 00 78 56 34 12"
)

length = int.from_bytes(data[6:8], "little")
ident = int.from_bytes(data[8:10], "little")
delta = int.from_bytes(data[10:12], "little", signed=True)
value = int.from_bytes(data[20:24], "little")

print(length, ident, delta, hex(value))

Expected output is 16 4660 -2 0x12345678.

OBSERVE

The initial hypotheses can now be checked against the evidence: ENUM is the intended signature, 0xA5 is the flags field, and the multi-byte numeric fields consistently use little-endian order. The length value 16 also matches the 16 bytes following the eight-byte header.

The bytes never changed. What changed was the interpretation attached to them.

EXPLAIN

Numbers and bytes provide values and bit patterns. Radix, signedness, byte order, text encoding and field structure are representation rules that give those patterns meaning.

Deliverable

Produce an annotated dump, field map, integer calculations, decoded flag bits, signed interpretation, text interpretation, byte-order argument, at least one rejected alternative hypothesis and a small verification program.