From Binary to Hexadecimal – Understanding How Computers Represent Numbers
Computers ultimately store bit patterns. A value only gains meaning when those bits are interpreted according to some representation.
This course teaches number systems from that perspective: not only how to convert values on paper, but how binary, hexadecimal, signed integers, masks, endianness and numeric formats appear in real software and hardware.
The course starts from first principles and assumes no prior knowledge of binary, hexadecimal, assembly language, or digital electronics.
You should be comfortable with basic integer arithmetic: addition, subtraction, multiplication, and division. When powers or other notation are needed, they are introduced before later material depends on them.
Programming and hardware examples are used to show where the representations appear in practice. You do not need prior C, Python, or assembly-language experience to begin the course.
By the end of the course you should be able to:
The number twelve can be written in several ways:
121100₂C₁₆14₈The value is the same. The representation is different.
This distinction is the foundation for the rest of the course.
In a positional numeral system, a digit’s value depends on both the digit and its position.
Decimal 347 means
3 × 10² + 4 × 10¹ + 7 × 10⁰.
Binary 1011₂ means
1 × 2³ + 0 × 2² + 1 × 2¹ + 1 × 2⁰ = 11.
The base tells us how many different digits are used before another position is needed.
| Base | Name | Digits |
|---|---|---|
| 2 | binary | 0–1 |
| 8 | octal | 0–7 |
| 10 | decimal | 0–9 |
| 16 | hexadecimal | 0–9, A–F |
Programming languages often use prefixes such as 0b1011
for binary and 0x2A for hexadecimal. Exact syntax varies
between languages and assemblers.
What does 10 mean in base 2, base 8 and base 16?
In any base b, the notation 10 has the
value 1 × b¹ + 0 × b⁰.
Therefore:
10₂ = 2₁₀10₈ = 8₁₀10₁₆ = 16₁₀The digit sequence alone is not enough. You must also know which base is being used.
Binary uses only 0 and 1. Each position has a power-of-two weight.
What decimal value does 10110₂ represent?
From right to left the weights are 1, 2, 4, 8, 16 and so on. Thus
10110₂ = 16 + 4 + 2 = 22₁₀.
A bit can hold 0 or 1. Four bits are often called a nibble. Eight bits normally form a byte. Wider machine values are described by an explicit bit width.
With n bits there are 2ⁿ possible patterns.
Four bits give 16 patterns; eight bits give 256.
A bit pattern does not have one inherent numeric meaning.
11111111₂ can be 255 as unsigned 8-bit, -1 as signed two’s
complement, or part of something that is not an integer at all. The
representation rule supplies the meaning.
For the prediction, 10110₂ = 16 + 4 + 2 = 22₁₀.
Binary makes widths and representation explicit: one more bit doubles the number of available bit patterns.
Binary is useful because digital hardware can reliably distinguish two states. The representation maps naturally to logic levels, registers and memory, but the bit pattern still needs context before it has a numeric meaning.
Convert 1111₂ and 10000000₂ to decimal.
Explain why eight bits provide 256 patterns but the largest unsigned
8-bit value is 255.
Hexadecimal uses 0–9 and A–F. One hexadecimal digit corresponds exactly to four binary bits.
What binary pattern is represented by A5₁₆?
A = 10, B = 11, through
F = 15. Therefore
A5₁₆ = 10×16 + 5 = 165₁₀.
Using four-bit groups:
A5₁₆ = 1010 0101₂.
The mapping is exact because four bits have 2⁴ = 16
possible patterns, exactly matching one hexadecimal digit.
Programming languages commonly use the 0xA5 notation.
Assemblers and debuggers may also use forms such as $A5 or
an h suffix.
The reverse conversion is equally direct: expand every hexadecimal
digit to four bits. For example, 3A₁₆ = 0011 1010₂.
The letters A–F are hexadecimal digits representing the values 10–15. They are not text characters merely because letters are used as symbols.
For the prediction, A₁₆ = 1010₂ and
5₁₆ = 0101₂, so A5₁₆ = 1010 0101₂.
Two hex digits represent one byte. This makes hex especially convenient for memory dumps, addresses, machine code and bit masks.
Hex is compact binary notation. Learning the 16 four-bit patterns removes much of the friction from low-level computing.
Convert FF₁₆ to binary and decimal. Explain why one hex
digit always maps to exactly four bits.
Octal uses digits 0–7. Each octal digit corresponds exactly to three binary bits.
Translate 111 101 010₂ one three-bit group at a
time.
Base-8 positions have weights 8⁰, 8¹,
8² and so on. For example, 572₈ = 378₁₀, and
101 111 010₂ = 572₈.
Octal remains visible in Unix-style permission notation such as
755, where each digit summarizes three permission bits.
For the prediction, 111₂ = 7₈, 101₂ = 5₈
and 010₂ = 2₈, so 111 101 010₂ = 752₈.
Three bits represent values 0–7, exactly the digit range of octal.
The value 755 in a Unix file mode is normally read as
octal when used as a numeric permission value. It does not mean decimal
seven hundred and fifty-five.
Octal does not change the number, only its representation. Its computing value comes from its direct three-bit grouping.
111111₂ to octal.17₈ to binary.Conversion changes representation while preserving value.
Are 1010₂, 12₈, 10₁₀ and
A₁₆ different numbers?
To convert to decimal, sum each digit times its positional weight:
2D₁₆ = 2×16 + 13 = 45₁₀.
To convert a decimal integer to another base, repeatedly divide by the target base and read the remainders in reverse.
Binary provides a convenient bridge. Group four bits for hexadecimal and three bits for octal:
0010 1101₂ = 2D₁₆
101 101₂ = 55₈
No: 1010₂ = 12₈ = 10₁₀ = A₁₆. The four notations
represent the same value in different bases.
The value remains unchanged throughout. Only the symbols and positional weights differ.
Conversion becomes much faster once common binary, octal and hexadecimal patterns are recognized directly.
Convert 42₁₀ to binary and hex, and FF₁₆ to
decimal.
Computing repeatedly uses quantities of the form 2ⁿ.
Knowing common powers of two makes addresses, bit masks and capacities
easier to read.
Here n tells us how many factors of 2 are multiplied
together. For example, 2³ = 2 × 2 × 2 = 8.
How many distinct values can eight bits represent?
Useful landmarks are:
| n | 2ⁿ |
|---|---|
| 0 | 1 |
| 1 | 2 |
| 2 | 4 |
| 3 | 8 |
| 4 | 16 |
| 8 | 256 |
| 10 | 1024 |
| 16 | 65536 |
| 20 | 1048576 |
| 32 | 4294967296 |
With n bits there are 2ⁿ bit patterns.
Eight bits therefore provide 256 patterns, representing 0 through 255
when interpreted as unsigned.
1 KiB = 1024 bytes, while the SI quantity
1 kB = 1000 bytes. MiB and GiB are corresponding binary
prefixes.
2⁸ = 256 means that eight bits provide 256 different
patterns. It does not mean that the largest unsigned value is 256.
Counting starts at zero, so the largest value is 255.
Eight bits form 2⁸ = 256 distinct bit patterns and can
therefore represent 256 distinct values once an interpretation is
chosen.
One hexadecimal digit is four bits, two hexadecimal digits are one byte, and four hexadecimal digits span 16 bits.
Powers of two connect bit width, ranges, memory sizes and address spaces.
An unsigned integer uses every bit to represent a non-negative magnitude.
What is the largest value that fits in eight bits?
With n bits there are 2ⁿ patterns, giving
the unsigned range 0 ... 2ⁿ − 1.
Examples:
| Width | Range |
|---|---|
| 8 bits | 0–255 |
| 16 bits | 0–65,535 |
| 24 bits | 0–16,777,215 |
| 32 bits | 0–4,294,967,295 |
11111111₂ = FF₁₆ = 255₁₀ when interpreted as unsigned
eight-bit data.
Fixed-width arithmetic cannot represent every integer. Under
modulo-2ⁿ arithmetic, keeping only eight bits makes
255 + 1 wrap to 0.
Eight bits provide 256 possible patterns, but the unsigned range is 0–255 because zero consumes one of those patterns.
For an unsigned 8-bit integer, the maximum is
2⁸ − 1 = 255, or 11111111₂ = FF₁₆.
The bits alone do not say “unsigned”. Their interpretation depends on the type and width supplied by context.
Width and interpretation matter when reading registers, protocols, file formats and machine code.
Find the range of a 12-bit unsigned integer and compute
250 + 10 with eight-bit wraparound.
Negative integers require an agreed interpretation of bit patterns.
Can 11111111₂ mean both 255 and −1?
Sign-magnitude reserves a sign bit and therefore has both +0 and −0.
Ones’ complement forms a negative value by inverting every bit and also has two zero encodings.
Two’s complement is the standard modern integer
representation. For n bits its range is
−2ⁿ⁻¹ ... 2ⁿ⁻¹ − 1. Eight bits therefore cover −128 through
127.
The top bit in two’s complement is not merely a separate minus flag that can be removed from the rest of the number. The entire bit pattern participates in the representation.
This is also why the range is asymmetric: an eight-bit two’s-complement integer has one more negative value than positive values.
The same byte FF₁₆ is 255 as unsigned eight-bit data and
−1 as signed eight-bit two’s complement.
Meaning requires context: width, signedness and encoding rule. Raw bits do not carry those labels themselves.
What is the signed 16-bit two’s-complement range? Explain the asymmetry around zero.
Two’s complement lets binary addition hardware operate naturally on positive and negative integers.
How can 11111011₂ represent −5?
To encode −5 in eight bits, write +5 as 00000101, invert
the bits to 11111010, then add one to obtain
11111011 or FB₁₆.
For an n-bit pattern whose top bit is one, another
interpretation method is unsigned-value minus 2ⁿ:
251 - 256 = -5.
00000101 + 11111011 = 1 00000000. Discarding the carry
beyond eight bits leaves zero.
Negation is invert-plus-one, but signed eight-bit −128 has no representable +128 counterpart. Code must account for this boundary case.
For the prediction, 11111011₂ is −5 in 8-bit two’s
complement: invert 00000101₂ to 11111010₂ and
add 1 to obtain 11111011₂.
Carry and signed overflow answer different questions, so CPUs may expose separate flags for them.
Two’s complement is naturally understood as modulo-2ⁿ
arithmetic plus a signed interpretation of the bit patterns.
Encode −1, −42 and −128 as eight-bit hexadecimal. Interpret
80₁₆ as both unsigned and signed.
Binary addition follows ordinary positional arithmetic with only two digits.
What is 1111₂ + 0001₂ in a four-bit register?
1+1=10₂, producing a carry. Thus
1111 + 0001 = 1 0000. Keeping four result bits gives zero
while the carry records unsigned overflow beyond the width.
Subtraction can use borrow directly or addition of a two’s-complement negative operand.
Many CPUs record properties of the result. Common flags include:
For signed eight-bit arithmetic, 7F₁₆ + 01₁₆ = 80₁₆. The
mathematical +128 cannot be represented, so CPUs with an overflow flag
report that condition.
Flag names and exact behavior are not identical across CPU architectures. The instruction-set documentation defines which flags an instruction changes and what they mean there.
Carry and signed overflow are also different conditions: an operation can set one without setting the other.
With only four bits, 1111₂ + 0001₂ = 1 0000₂. The
register retains the low four bits 0000₂, while the fifth
bit is the carry out.
Carry and signed overflow answer different questions about the same bit-level result.
Flags allow later instructions to branch, compare and extend arithmetic without recomputing the operation.
Compute FF₁₆ + 01₁₆ and discuss C, Z and V for eight-bit
arithmetic.
Bitwise operations process corresponding bits independently.
What is 1010 AND 1100?
AND requires both bits to be one. OR requires at least one. XOR is one when the bits differ. NOT inverts every bit.
For A=1010₂, B=1100₂: AND gives
1000, OR 1110, XOR 0110, and
four-bit NOT A gives 0101.
Logical shifts move bits and fill with zero. Arithmetic right shift commonly preserves the sign bit. Rotates wrap bits around instead of discarding them; rotate-through-carry also includes the carry flag.
For the prediction, compare corresponding bits:
1010 AND 1100 = 1000.
Width is essential: NOT zero has a different numerical result at 8 and 16 bits.
These operations underpin register manipulation, protocols, graphics, compression and many low-level algorithms.
Compute AND, OR and XOR for 3C₁₆ and
0F₁₆.
A bit mask selects particular bits while leaving neighboring bits available for other meanings.
How can bit 0 be set in 10100100₂ without changing the
other bits?
With MASK=00000001₂: set using OR, clear using AND with
NOT MASK, toggle using XOR, and test using AND.
Fields spanning several bits can be isolated with a mask and shifted down before interpretation.
Suppose bit 7 means ENABLE, bit 3 IRQ and bit 0 READY. Then
10001001₂ carries three independent Boolean flags in one
byte.
Set bit 0 by OR-ing with 00000001₂:
10100100₂ OR 00000001₂ = 10100101₂. The other bits are
unchanged.
Hex makes masks compact: 11110000₂ = F0₁₆, while
00001111₂ = 0F₁₆.
Bit fields efficiently pack small values, and masks let software manipulate one field without damaging the rest.
Create a mask selecting bits 2 and 5, then show how to toggle bit 5
in 34₁₆.
Memory can be modeled as a sequence of addressable bytes. An address identifies a location; a value may occupy one or several consecutive bytes.
If a 32-bit value starts at address 0x1000, which byte
addresses does it occupy?
In byte-addressable memory the address normally increases by one per
byte: 0x1000, 0x1001, 0x1002,
0x1003. A 32-bit value occupies four bytes.
An offset is a displacement relative to a base
address. Base 0x2000 plus offset 0x34 gives
0x2034.
With n address bits, 2ⁿ address patterns
exist. On a byte-addressed system, 16 address bits can identify 65,536
byte locations.
Architectures may prefer or require multi-byte values to begin at particular address boundaries. A 32-bit value may be naturally aligned at an address divisible by four.
The predicted 32-bit value occupies four bytes: 0x1000,
0x1001, 0x1002 and 0x1003.
Hexadecimal maps cleanly to address bits: a 16-bit address is four
hex digits from 0000 through FFFF.
Debugger and hex-editor displays require you to distinguish the address of data from the value stored there.
Which addresses are occupied by a 16-byte block beginning at
0x3FF0? Compute base 0x8000 plus offset
0x2A.
Endianness describes the byte order used to store a multi-byte value. It does not change the numerical value itself.
How can 0x12345678 occupy four bytes beginning at
0x1000?
Big-endian stores the most significant byte first:
12 34 56 78 at increasing addresses.
Little-endian stores the least significant byte first:
78 56 34 12.
The bit notation inside each normally written byte is not reversed. Here, endianness is about ordering bytes of a multi-byte value.
The Motorola 68000 family is associated with big-endian storage, while x86 uses little-endian. Protocols and file formats can define byte order independently of the host CPU.
A dump showing 78 56 34 12 can encode the 32-bit number
0x12345678 when the format specifies little-endian. Raw
bytes need width, type and byte-order context.
For the prediction, 0x12345678 may occupy memory as
12 34 56 78 in big-endian or 78 56 34 12 in
little-endian order from the lowest address. Byte order is part of the
representation.
A hex dump normally displays bytes in increasing memory-address order, which need not match the written digit order of a multi-byte number.
Endianness matters when exchanging data among CPUs, networks, binary files, emulators and debugging tools.
Write 0xA1B2C3D4 as four bytes in both big- and
little-endian order.
BCD encodes decimal digits individually rather than representing the entire number as one binary integer.
How can decimal 42 be stored when each decimal digit receives four bits?
Packed BCD gives each digit one nibble:
42 -> 0100 0010₂ -> 0x42. Ordinary binary 42 is
instead 00101010₂ = 0x2A.
Only nibble values 0–9 are normally valid decimal digits. Unpacked BCD commonly gives each digit a whole byte.
BCD has been useful in calculators, clocks, counters, financial systems and historical computers where decimal digits need direct preservation.
With four bits per decimal digit, 4 is 0100 and 2 is
0010. Packed BCD for 42 is therefore
0100 0010₂ = 0x42.
A byte 0x42 can mean ordinary integer 66 or packed-BCD
decimal 42. The format supplies the meaning.
BCD reinforces a central theme: bits need a representation rule before they become a number.
Encode 1987 as packed BCD and state its storage size.
Fixed-point represents fractional values using an implied, fixed binary-point position.
With four integer bits and four fractional bits, what does
00111000₂ represent?
Interpreted as 0011.1000₂, the value is 3.5. The stored
integer is 56 and four fractional bits imply division by
2⁴ = 16.
With f fractional bits, the resolution is
1 / 2^f.
Two’s-complement signed integers can use the same scaling rule for signed fixed-point values.
Fixed-point offers predictable resolution and efficient arithmetic on systems without fast floating-point hardware, but software must manage scaling, range and overflow explicitly.
With four fractional bits, place the binary point after the upper
four bits: 0011.1000₂ = 3 + 1/2 = 3.5.
No physical binary-point symbol is stored. Its location belongs to the format definition.
Fixed-point is scaled integer arithmetic, directly connecting fractional numbers to integer width and overflow.
What is the resolution with eight fractional bits? Interpret stored integer 384 using that scale.
Floating-point stores a significand together with an exponent so the binary point can effectively move.
How can the same 32-bit format cover both very large and very small magnitudes?
IEEE 754 binary32 contains one sign bit, eight exponent bits and 23
explicit fraction bits. Normal values conceptually follow
(-1)^sign × significand × 2^exponent; the exponent uses a
bias and the leading one of a normal binary significand is implicit.
IEEE 754 also defines signed zeros, infinities, NaNs and subnormal values.
1.0 = 0x3F800000The binary32 pattern 0x3F800000 is:
0 01111111 00000000000000000000000
127 − 127 = 01.0₂The value is therefore (+1) × 1.0₂ × 2⁰ = 1.0.
A fraction can be converted by repeatedly multiplying its fractional part by 2 and recording the integer part:
0.5 × 2 = 1.0 → the first bit is 1, so
0.5₁₀ = 0.1₂0.25 × 2 = 0.5, then 0.5 × 2 = 1.0 →
0.25₁₀ = 0.01₂For 0.1₁₀, the process does not terminate: the bit
pattern repeats. There is therefore no finite binary fraction exactly
equal to one tenth.
One tenth has no finite binary fractional expansion, just as one
third has no finite decimal expansion. It must be rounded to a
representable binary floating-point value. Arithmetic on rounded
operands can therefore differ slightly from ideal decimal arithmetic. In
ordinary binary64 arithmetic, for example, 0.1 + 0.2 is
typically represented as approximately 0.30000000000000004,
not exactly 0.3.
Exponent bits provide range while significand bits provide precision. Floating-point is a controlled compromise, not arbitrary-precision real arithmetic.
0x3F800000 demonstrates concretely how the sign,
exponent and fraction fields combine to produce 1.0. The 0.1 example
also shows why not every decimal fraction can be stored exactly.
A floating-point bit pattern cannot be interpreted as an ordinary integer while preserving its numerical meaning; the field structure defines the representation.
IEEE 754 standardizes behavior, but programs still need to account for rounding, special values and comparisons.
Why is exact floating-point equality often more subtle than integer equality?
Bit patterns represent much more than ordinary integers. This chapter introduces Gray code, biased/excess notation, character codes and packed data.
Must the bit pattern 01000001 always mean the number
65?
In binary-reflected Gray code, neighboring values differ in exactly one bit. This is useful when multiple bits should not transition simultaneously, such as in some position encoders.
A binary value b can be converted with
g = b XOR (b >> 1).
An exponent can be stored as an unsigned field plus an agreed bias. With bias 127, stored value 130 represents exponent (130-127=3).
IEEE 754 binary32 uses this principle for its exponent field, with reserved field values receiving special meanings.
ASCII assigns A the code value 65, or 0x41.
Unicode defines code points, such as U+0041 for A. An
encoding such as UTF-8 then determines the actual stored bytes.
A code point and its byte encoding are not the same concept.
One byte can contain several bit fields. A hypothetical status register might use bit 7 for ready, bit 6 for error, bits 5–4 for a mode and bits 3–0 for a counter.
No. 01000001₂ = 0x41 is unsigned integer 65, but the
same pattern can, for example, represent the ASCII character
A.
0x41 may be unsigned 65, signed 65, ASCII A, part of an
instruction or a field inside a larger structure. Context supplies
meaning.
Computers store bits. Types, protocols, instruction sets and file formats supply their semantics.
Compute the Gray code for binary 1010. What exponent
does stored value 124 represent with excess-127?
Assembly exposes the distinction among value, syntax, bit width and machine meaning. A number may be an immediate constant, address, mask, offset or part of an instruction encoding.
Do $10, #10, 0x10 and
10h always mean the same thing?
Assembler dialects use different literal conventions. Common forms
include decimal 42, hexadecimal 0x2A,
$2A in many classic assemblers, 2Ah in some
Intel-style contexts, and forms such as %101010 for
binary.
Always interpret syntax in the context of the actual assembler.
A numeric operand can be the value itself or identify a location from which a value is read.
On 6502-style syntax, LDA #$2A commonly loads immediate
value 0x2A, while LDA $2A accesses memory
according to the instruction’s addressing mode.
On Motorola 68000, MOVE.B #$2A,D0 uses an immediate
value, while MOVE.B $002A,D0 refers to memory. Size
suffixes such as .B, .W and .L
make operand width visible in many 68k dialects.
AVR uses its own register and instruction constraints. For example,
LDI R16, 0x2A loads a constant into a permitted register. A
source literal still has to fit the operand range and instruction
encoding.
x86 has multiple major assembler dialects. Intel and AT&T syntax can express the same machine operation differently, so no single textual convention should be mistaken for universal assembly syntax.
Registers and instruction fields have finite widths. Bit masks such
as 0x0F are commonly used to select fields.
A disassembler interprets raw bytes as instructions for a chosen ISA and address. Bytes that are actually data can look like plausible instructions when context is wrong.
No. Prefixes and suffixes are assembler syntax, and the meanings of `# Numbers in assembly
Assembly exposes the distinction among value, syntax, bit width and machine meaning. A number may be an immediate constant, address, mask, offset or part of an instruction encoding.
Do $10, #10, 0x10 and
10h always mean the same thing?
Assembler dialects use different literal conventions. Common forms
include decimal 42, hexadecimal 0x2A,
$2A in many classic assemblers, 2Ah in some
Intel-style contexts, and forms such as %101010 for
binary.
Always interpret syntax in the context of the actual assembler.
A numeric operand can be the value itself or identify a location from which a value is read.
On 6502-style syntax, LDA #$2A commonly loads immediate
value 0x2A, while LDA $2A accesses memory
according to the instruction’s addressing mode.
On Motorola 68000, MOVE.B #$2A,D0 uses an immediate
value, while MOVE.B $002A,D0 refers to memory. Size
suffixes such as .B, .W and .L
make operand width visible in many 68k dialects.
AVR uses its own register and instruction constraints. For example,
LDI R16, 0x2A loads a constant into a permitted register. A
source literal still has to fit the operand range and instruction
encoding.
x86 has multiple major assembler dialects. Intel and AT&T syntax can express the same machine operation differently, so no single textual convention should be mistaken for universal assembly syntax.
Registers and instruction fields have finite widths. Bit masks such
as 0x0F are commonly used to select fields.
A disassembler interprets raw bytes as instructions for a chosen ISA and address. Bytes that are actually data can look like plausible instructions when context is wrong.
, #, 0x and h depend on the
language and context. The syntax rules of the particular assembler are
authoritative.
Assembly source contains human-readable symbols and numeric notation. The CPU receives encoded instruction bit fields.
Number representation is fundamental to assembly because operands, addresses, registers, masks, offsets and encodings are all finite bit fields.
Why are #$2A and $2A different in many
classic assembler dialects? Why must you know the dialect before
interpreting 10h?
C and Python can express the same binary and hexadecimal values, but their integer models differ substantially.
Will 255 + 1 always behave the same way in C and
Python?
Hexadecimal and decimal literals such as 0x2A and
42 work the same way in both languages. Binary literals
such as 0b101010 are supported in Python and, from C23, in
C (many compilers, including GCC and Clang, have long accepted them as
an extension). In C, language type and literal rules determine the
resulting type. Python int is not confined to a fixed 8-,
16-, 32- or 64-bit width.
When machine width matters, <stdint.h> provides
types such as uint8_t, uint16_t and
int32_t when the implementation has suitable exact-width
integer types.
Unsigned C arithmetic wraps modulo 2ⁿ for the type
width. Do not assume the same rule for signed overflow; signed integer
overflow is not defined by C as ordinary two’s-complement
wraparound.
Python integers grow as required within practical resource limits. To model an 8-bit register explicitly, impose a width:
x = (255 + 1) & 0xFFThe result is zero.
Bitwise operations look familiar in both languages, but type widths, promotions and negative-value behavior require language-specific care.
Python can parse bases with int(text, base) and format
values with forms such as f"{42:08b}" and
f"{42:02X}".
C offers conversion routines such as strtoul; portable
formatting of fixed-width integers can use the macros in
<inttypes.h>.
No. The result depends on type and language rules. Python integers normally grow as needed, while arithmetic involving bounded C integer types follows different rules. Both the type and language semantics matter.
The source literal 0xFF is a numeric value. Storing or
interpreting it as an 8-bit signed value, 8-bit unsigned value or wider
type is a separate operation.
Programming languages add rules on top of bit representations. Reliable low-level code requires understanding both the machine model and the language model.
Why does (255 + 1) & 0xFF produce zero in Python?
Why should signed C overflow not be assumed to work the same way?
Digital electronics connects bit patterns to physical signals, registers, measurements and protocols. Numbers do not merely describe mathematics: they control pins and encode observations of the physical world.
If an 8-bit GPIO register contains 0x81, which bits are
set?
Logical zero and one are abstractions. Physical circuits use voltage ranges interpreted as LOW and HIGH; exact thresholds come from the component datasheet.
Signals can be active-low, meaning the asserted function corresponds to an electrically LOW line. Naming conventions vary.
For PORT = 10100001₂ = 0xA1, individual bits may control
separate functions.
Typical software operations include setting a bit with
PORT |= (1 << 3), clearing it with
PORT &= ~(1 << 3), and toggling it with XOR.
Real hardware must be checked against its datasheet. Some registers provide dedicated SET/CLEAR aliases, write-one-to-clear flags or other semantics for which ordinary read-modify-write is inappropriate.
Register documentation commonly specifies addresses or offsets, bit positions, field widths, reset values, access properties and meanings of encoded field values.
A value such as 0x82 is only meaningful once its
register and field definitions are known.
An idealized N-bit ADC provides 2^N digital codes. A
10-bit converter has 1024 codes, usually numbered 0 through 1023.
A simple idealized unipolar conversion can be approximated by
V ≈ code / (2^N - 1) × Vref
but the exact transfer function, reference, tolerances and endpoint conventions come from the datasheet. The formula above is therefore a learning model, not a universal ADC law.
A DAC maps digital codes in the opposite direction toward an analog output; its width likewise determines the number of available codes.
Logic analyzers capture digital samples and may display bits, bytes, hexadecimal values, timestamps or decoded UART/SPI/I²C fields.
The byte 10100101₂ = 0xA5 might be a command, address,
flags or payload. Protocol context supplies meaning. Serial analysis may
additionally require bit order, clocking, framing and multi-byte byte
order.
For the prediction, 0x81 = 1000 0001₂, so bits 7 and 0
are set.
Hardware registers and protocols demonstrate the course’s central idea: bit pattern plus representation rules produces meaning.
Reading a datasheet constantly requires translating among bit positions, masks, hexadecimal values, numerical quantities and physical behavior.
Which bits are set in 0x81? How many codes does a 12-bit
ADC have? Why can ordinary read-modify-write be wrong for some status
registers? Why must an ADC conversion formula be checked against the
datasheet?
Network packets and binary files are sequences of bytes. Understanding them requires knowing which bytes form fields, the byte order of multi-byte values and the meaning assigned by the format.
The bytes 12 34 form a 16-bit integer. Is the value
0x1234 or 0x3412? Write down your hypothesis
before continuing.
Many multi-byte fields in Internet protocols use network byte
order, which is big-endian. A 16-bit 0x1234 field
is therefore transmitted as 12 34.
This does not require the host CPU itself to be big-endian. Protocol representation and native machine representation are separate concerns.
IPv4 is 32 bits. The familiar 192.168.1.10 corresponds
to bytes C0 A8 01 0A.
A common 48-bit MAC address can be displayed as six hexadecimal
bytes, for example 02:12:34:56:78:9A. Hexadecimal maps
naturally to byte-oriented representations.
Packets may contain versions, lengths, types, flags, ports, sequence numbers, checksums and payloads. Some are bit fields; others are multi-byte integers. The protocol specification defines width, byte order and semantics.
Binary formats often begin with characteristic signatures. For
example, 89 50 4E 47 0D 0A 1A 0A is the eight-byte
signature at the start of a PNG file.
A signature can aid identification, but does not by itself prove that the rest of a file is structurally valid.
Consider a fictional record:
01 03 12 34 41 42 43
Its specification says byte 0 is a version, byte 1 a payload length, bytes 2–3 a big-endian ID and the remainder payload.
The record therefore contains version 1, length 3, ID
0x1234, and payload 41 42 43, which can be
interpreted as ASCII ABC.
Without the specification these are simply seven bytes.
Binary parsers move through structures using field offsets and lengths. A field beginning at offset 4 with length 2 occupies bytes 4 and 5; the next field begins at offset 6.
An incorrect width or offset shifts subsequent interpretation.
There is no single answer to the prediction without a byte-order
rule: big-endian gives 0x1234, while little-endian gives
0x3412.
The same byte can be part of an IP address, character, flags, integer or opaque payload.
A hex dump exposes stored or transmitted bytes. A format or protocol specification tells us how to group and interpret them.
How is 0xBEEF stored as a big-endian 16-bit field? Which
bytes encode IPv4 127.0.0.1? Why is a magic number
insufficient to validate an entire file?
When software crashes, a file is undocumented or an old system lacks specifications, raw bytes may appear before their structure is known. Number representation becomes an analysis tool.
You find the bytes 48 65 6C 6C 6F 00. Are they integers,
machine code or text?
Separate observations from hypotheses. A hex editor or memory dump may show offsets or addresses, hexadecimal bytes, a text column and highlighted regions.
Readable characters are evidence worth investigating, not proof that a field is text.
Useful observations include repeated zeros, recurring values, text-like sequences, possible lengths or offsets, known signatures, regular blocks and bit fields whose bits change predictably.
A hypothesis becomes stronger when it explains multiple independent examples.
Bytes 34 12 may represent big-endian 0x3412
or little-endian 0x1234.
If the field is hypothesized to be a length, test whether that value agrees with the surrounding structure.
Debugger addresses are commonly hexadecimal. Address differences
reveal sizes: 0x1040 - 0x1000 = 0x40 = 64.
A disassembler interprets bytes as instructions for a selected CPU and starting point. That does not prove those bytes are executable code.
Data can decode into syntactically valid but meaningless instructions. Ask whether the CPU and start address are known, whether control flow reaches the region, whether the sequence is plausible and whether the bytes might instead be data.
Code, tables, strings and constants may live close together.
41 42 43 44 could be ASCII ABCD, four
integers, portions of larger values or instruction bytes. Context
determines the useful interpretation.
Consider the fictional bytes
03 00 08 00 41 42 43 00.
One hypothesis is type 3, flags 0, a little-endian length of 8 and four data bytes. This is a model, not yet a fact. Comparing additional records can test whether those fields behave consistently.
With systems and test data you control, changing one known value at a time and comparing before and after is powerful.
If enabling a flag changes exactly one bit, that is useful evidence. If changing a counter tracks one field, its representation becomes clearer.
Several interpretations are possible without context. As ASCII/UTF-8,
48 65 6C 6C 6F 00 gives Hello followed by a
zero byte; that is a strong hypothesis, not proof by itself.
Data reverse engineering is often a process of constructing hypotheses that can be weakened or strengthened rather than instantly guessing the correct format.
Bases, signedness, byte order, bit fields, addresses and text encodings make raw bytes easier to reason about. Context and repeated observations distinguish plausible interpretation from documented structure.
Why can data produce apparently valid disassembly? What could
00 01 mean under different byte orders? Why are multiple
examples more useful than one dump when reconstructing a structure?
Binary, octal, decimal and hexadecimal dominate modern computing, but positional notation can use many other bases. Exploring them reinforces that a base is a representation rule rather than a property of the number itself.
What do you think 10 means in base 3? Write down your
answer before continuing.
Base 3 uses digits 0, 1 and 2. For example,
102₃ = 1 × 9 + 0 × 3 + 2 = 11₁₀.
Ternary systems demonstrate that digital representation is not mathematically restricted to exactly two symbols. Ternary logic has also appeared in computing research and historical machines.
Base 12 requires digit values for ten and eleven in addition to 0–9.
In this course we use A for ten and B for
eleven, so A₁₂ = 10₁₀ and B₁₂ = 11₁₀. Twelve
has several divisors — 2, 3, 4 and 6 — so some fractions are
compact.
One half in base 12 is 0.6₁₂, because 6/12 = 1/2.
Vigesimal systems use base 20 and occur historically in multiple
languages and cultures. For representation, the key point is simple:
10₂₀ means twenty rather than ten.
Base 36 can use 0–9 and A–Z as its 36 digit values. Thus
Z₃₆ = 35₁₀ and 10₃₆ = 36₁₀.
It can provide compact textual forms for non-negative integers, although a format must define its alphabet and case rules.
Sexagesimal systems have ancient roots. Base-60 structure remains visible in time and angle measurement: 60 seconds per minute, 60 minutes per hour, 60 arcminutes per degree and 60 arcseconds per arcminute.
These conventions are not all simply modern positional base-60 notation, but they illustrate how a numerical subdivision can persist.
Factors of the base affect which fractions terminate.
In base 10, 1/2 and 1/5 terminate while 1/3 repeats. In base 2, 1/2 terminates but decimal 1/10 does not have a finite binary fractional representation. This connects directly to floating-point behavior.
A base and a storage width are separate concepts. A value displayed as base-36 text may be stored internally as an ordinary binary integer. A byte can likewise be displayed in decimal, hexadecimal or another base without changing its value.
Computing history includes representations unlike today’s most familiar binary conventions. Some machines emphasized decimal arithmetic or used unusual word sizes and character encodings.
Historical data must therefore be interpreted according to the documented representation of the relevant machine rather than modern assumptions.
For the prediction, 10₃ means three:
1 × 3¹ + 0 × 3⁰ = 3.
The notation 10 can represent 2, 3, 8, 10, 12, 16, 20,
36, 60 or another value depending on the base.
A positional system defines digits, positional weights and a radix. The number is the abstract value; the notation is its representation.
What is 10₁₂ in decimal? What is Z₃₆? Why
can base 12 express some common fractions more compactly than base
10?
This project combines the ideas from EduNumbers. You will use representation, byte order, signedness, bit fields, text and structure to explain a synthetic unknown buffer.
0000: 45 4E 55 4D 01 A5 10 00 34 12 FE FF 50 4C 4F 4F
0010: 53 00 00 00 78 56 34 12
You initially know only that the format was created for this exercise, every field starts on a byte boundary, and the buffer is complete.
Before calculating, identify regions that look like text, flags or possible little-endian integers. Record hypotheses first.
The first four bytes decode as ASCII ENUM, making them a
strong signature candidate.
Byte 0x01 at offset 0x04 is plausible as a
version. Byte 0xA5 at 0x05 is
10100101₂, with bits 7, 5, 2 and 0 set.
Bytes 10 00 become 16 as little-endian 16-bit; as
big-endian they would be 4096. Since the complete buffer contains 16
bytes after its eight-byte header, 16 is a plausible length, while
34 12 becomes 0x1234 = 4660.
Bytes FE FF become 65534 unsigned or -2 as signed 16-bit
two’s complement. The bits alone do not specify signedness.
Bytes 50 4C 4F 4F 53 00 00 00 resemble an eight-byte
null-padded ASCII field containing PLOOS.
Finally, 78 56 34 12 becomes 0x12345678 as
little-endian. Multiple consistent fields strengthen the little-endian
hypothesis.
| Offset | Size | Possible field | Interpretation |
|---|---|---|---|
| 0x00 | 4 | magic | ASCII ENUM |
| 0x04 | 1 | version | 1 |
| 0x05 | 1 | flags | 0xA5 |
| 0x06 | 2 | length | 16, little-endian |
| 0x08 | 2 | id | 0x1234 |
| 0x0A | 2 | delta | -2 signed |
| 0x0C | 8 | name | PLOOS, null-padded |
| 0x14 | 4 | value | 0x12345678 |
For this synthetic exercise that is the intended structure. In real reverse engineering, more samples or documentation would be required before treating such a model as established fact.
data = bytes.fromhex(
"45 4E 55 4D 01 A5 10 00 34 12 FE FF "
"50 4C 4F 4F 53 00 00 00 78 56 34 12"
)
length = int.from_bytes(data[6:8], "little")
ident = int.from_bytes(data[8:10], "little")
delta = int.from_bytes(data[10:12], "little", signed=True)
value = int.from_bytes(data[20:24], "little")
print(length, ident, delta, hex(value))Expected output is 16 4660 -2 0x12345678.
The initial hypotheses can now be checked against the evidence:
ENUM is the intended signature, 0xA5 is the
flags field, and the multi-byte numeric fields consistently use
little-endian order. The length value 16 also matches the 16 bytes
following the eight-byte header.
The bytes never changed. What changed was the interpretation attached to them.
Numbers and bytes provide values and bit patterns. Radix, signedness, byte order, text encoding and field structure are representation rules that give those patterns meaning.
Produce an annotated dump, field map, integer calculations, decoded flag bits, signed interpretation, text interpretation, byte-order argument, at least one rejected alternative hypothesis and a small verification program.