What Windows reads before running a PE
A PE file's DOS header, NT headers, data directories, and section table explain where code lives on disk and in memory, and where malware analysis should begin.
Introduction
A Windows executable is not just a stream of instructions. Its PE headers tell the system how to map the file into memory, where execution starts, and which other files or data structures it needs. They also give an analyst a useful first view of a suspicious sample before any code runs.
PE stands for Portable Executable. The same format underlies executables and DLLs, with PE32 and PE32+ variants. This post follows the headers of an image file and uses a small address calculation to connect bytes on disk with addresses seen in a debugger.
Finding the NT headers
A PE image starts with an MS-DOS header. Its first two bytes are MZ (0x4D5A). The field at file offset 0x3C, called e_lfanew, contains the file offset of the PE signature. At that location, the four bytes should read PE\0\0.
File offset
0x0000 MS-DOS header (MZ)
0x003C e_lfanew -> offset of PE signature
... DOS stub / padding
e_lfanew PE\0\0 | COFF file header | Optional Header | Section TableThe DOS stub exists for compatibility; its common message about DOS mode is not part of the code Windows starts executing. e_lfanew should be checked against the file size before following it. A parser that trusts offsets from an untrusted sample can read beyond the file.
The NT headers do not have a fixed file offset. If e_lfanew is 0x80, the PE signature starts at 0x80; if it is 0xF0, it starts at 0xF0. This field is a little-endian 32-bit integer. In a hex editor, the bytes 80 00 00 00 at 0x3C represent 0x80.
The signature is followed by a 20-byte COFF file header. The Optional Header comes next, and the COFF header declares its size. The Section Table starts after that block. Hard-coding an Optional Header size is a poor way to find the section entries: PE32 and PE32+ differ, and SizeOfOptionalHeader provides the correct boundary.
COFF header and Optional Header
The PE signature is followed by the COFF file header. Machine identifies the target architecture, NumberOfSections tells us how many section entries follow, and SizeOfOptionalHeader determines where the section table begins. Characteristics includes flags such as whether the image is an executable or DLL.
For an image, the so-called Optional Header is required. Its Magic value distinguishes PE32 (0x10B) from PE32+ (0x20B). This distinction changes the size of some fields, including ImageBase. It does not mean every field becomes 64 bits wide.
Machine and Magic answer different questions. Machine identifies the target architecture; Magic identifies the Optional Header variant. A .exe or .dll extension cannot provide either answer. Nor should the COFF TimeDateStamp be treated as a verified build date: it is a value stored in the file and can be changed.
Several fields are especially useful when reading a sample:
| Field | What it tells us |
|---|---|
AddressOfEntryPoint | RVA where execution begins after the image is loaded. |
ImageBase | Preferred base address for the mapped image. |
SectionAlignment / FileAlignment | Alignment of sections in memory and on disk. |
SizeOfImage | Space reserved for the image in memory, including aligned sections. |
SizeOfHeaders | Size of the headers as laid out in the file. |
Subsystem | Intended execution environment, such as console or GUI. |
An entry point in an unusual section is a reason to inspect that section, not a verdict on its own. Packers often start in a small unpacking stub, but legitimate protectors and unusual build pipelines can produce similar layouts.
AddressOfEntryPoint is not an absolute virtual address. If it is 0x1234 and the image loads at 0x140000000, execution starts at 0x140001234. The actual base can differ from the preferred ImageBase, including because of ASLR. The RVA remains relative to the beginning of the image, which makes it useful when moving between the file and a debugger.
Data Directories
The Optional Header ends with Data Directories. Each entry gives an address and size for a structure such as imports, exports, resources, relocations, or TLS. Most addresses here are RVAs. The certificate table is an exception: its address is a file offset because certificates are not mapped into the image in the usual way.
The Import Directory is a good place to see which DLLs and functions the image declares. It is only a partial picture of behavior: code can resolve APIs at runtime or load additional modules. Likewise, a sparse import table alone does not establish that the file is packed.
When parsing directories, respect both NumberOfRvaAndSizes and SizeOfOptionalHeader. Do not assume a directory starts at a section boundary or lives in a section with a particular name.
Imports: from a DLL name to the IAT
The Import Directory points to an array of descriptors. Each descriptor represents a DLL and includes references to its name and imported functions. The Import Lookup Table describes functions by name or ordinal. The Import Address Table, or IAT, is used by the code to access resolved addresses. Before resolution it can contain information similar to the lookup table; during loading, function addresses are prepared there.
This is why finding a string such as CreateFileW in the file can help but cannot settle what the code does. The name might be an import, unused text, or absent because the function is resolved at runtime. To understand a particular call, follow its reference from the code to the IAT and inspect the target after loading.
Relocations and TLS
If the image cannot load at its preferred ImageBase, some absolute references need adjustment. The base relocation directory identifies locations that receive the difference between the preferred and actual base. The loader does not rewrite every RVA in the file; it adjusts specific address-dependent values.
The TLS directory deserves a separate look. It can contain callbacks invoked during process or thread initialization before the usual entry point. Beginning a debugging session directly at AddressOfEntryPoint can miss that code. TLS callbacks are also a normal feature of the format, so their presence alone says nothing about intent.
Sections and the disk-to-memory mapping
Each section-table entry describes two views of the same region. PointerToRawData and SizeOfRawData locate its initialized bytes in the file. VirtualAddress and VirtualSize describe its position and size in the mapped image. Characteristics includes permissions such as readable, writable, and executable.
Consider a section with VirtualAddress = 0x1000 and PointerToRawData = 0x400. An entry point RVA of 0x1234 lies 0x234 bytes into that section. If those bytes are present in the raw section data, their file offset is:
file offset = RVA - VirtualAddress + PointerToRawData
= 0x1234 - 0x1000 + 0x400
= 0x634At a loaded base of 0x140000000, the same RVA refers to virtual address 0x140001234. An RVA is relative to the loaded image base; a file offset is a position in the file. Confusing the two leads to reading the wrong bytes.
The conversion only works for data actually present in that section’s raw range. If VirtualSize exceeds SizeOfRawData, the remaining mapped space is zero-filled. Headers also have their own mapping, so a general-purpose parser needs bounds checks rather than applying the section formula blindly.
The reverse can happen too: SizeOfRawData may exceed VirtualSize because FileAlignment adds padding. Those trailing bytes are not automatically meaningful content. First find the section containing the RVA, then verify that the calculated position corresponds to bytes actually present in the file.
In a typical PE, .text holds code, .rdata read-only data, .data writable data, and .rsrc resources. These names are conventions rather than types enforced by the loader. A randomly named section can contain valid code, while a section named .text may have an unusual layout.
A complete pass through the numbers
Consider a test PE32+ with e_lfanew = 0x80, two sections, and the following values. They are invented to make the calculations clear; they do not describe a real sample.
| Structure or field | Value |
|---|---|
| PE signature | File offset 0x80 |
AddressOfEntryPoint | RVA 0x1234 |
ImageBase | 0x140000000 |
.text | RVA 0x1000, raw 0x400, virtual/raw sizes 0x600 / 0x600 |
.rdata | RVA 0x2000, raw 0xA00, virtual/raw sizes 0x400 / 0x400 |
| Import Directory | RVA 0x2100, size 0x80 |
Look for the signature at 0x80, not immediately after the DOS Header. The entry point falls in .text because 0x1234 is between 0x1000 and 0x1600. Its file offset is 0x634, as calculated earlier. If the image loads at its preferred base, its virtual address is 0x140001234.
The Import Directory falls in .rdata: 0x2100 - 0x2000 = 0x100 bytes from the start of the section. In the file it begins at 0xA00 + 0x100 = 0xB00. Its declared size, 0x80, extends the range to 0xB80, still within .rdata’s 0x400 raw bytes. Checking the complete range matters: a valid start address does not guarantee that the full structure is present.
If the process loads at 0x150000000, the entry point moves to 0x150001234; file offsets 0x634 and 0xB00 do not change. The base difference is 0x10000000, and relocations adjust absolute values that depend on that base. The Import Directory is still found through RVA 0x2100 within the image.
Reading a suspicious sample
The headers help form testable questions. Does the entry point land in a section with little raw data? Is a section both writable and executable? Is SizeOfRawData much smaller than VirtualSize? Are imports unexpectedly sparse for the behavior seen at runtime? Does a directory point outside the file or overlap another structure?
These are triage signals, not malware rules. A single field rarely proves intent. Compare the header layout with the bytes it references, then check the resulting hypothesis in a disassembler or debugger. A packed sample, for example, may begin in a stub that reconstructs code and imports before transferring control to its original entry point.
A concrete pass makes those checks easier to follow. Suppose a file has three sections: .text, .rdata, and UPX1. Its entry point falls in executable UPX1, while the import table lists few functions. This combination suggests inspecting the code in UPX1 first. It does not yet prove what the file does, or even that it is malicious.
Calculate the entry point’s file offset and inspect those bytes. Then check whether the code writes to another memory region, changes permissions, or transfers control elsewhere. If a second entry point appears after unpacking, reconstruct the import view there and compare the image in memory with the original file. The UPX1 name does not replace that investigation: section names can be changed.
Two other limits matter. An unusual TimeDateStamp can hint at how the file was built, but it is not a reliable timeline. A certificate in the Certificate Table requires cryptographic validation and a trust decision; the directory’s presence does not authenticate the sample.
The useful habit is to move between three coordinates deliberately: file offset for bytes on disk, RVA for locations inside the image, and virtual address for the loaded process. Once those are separated, the rest of the PE structure becomes much easier to inspect.
- 01 Microsoft - PE Format learn.microsoft.com/en-us/windows/win32/debug/pe-format ↗