System registers
Reached through MRS and MSR, each system register is encoded by five op fields. Which of them you can read and write depends on your exception level, so most of these pages show an access table before anything else.
The four condition flags, packed into the top four bits of PSTATE. Most data-processing instructions write these as a side effect, and every conditional instruction reads them to decide whether to execute. Reachable directly through MRS and MSR, which is how software saves and restores flags.
The four interrupt mask bits of PSTATE, gathered into one register so they can be read and written in a single instruction. A set bit means the corresponding exception is masked at the current level.
Reports which exception level the processor is currently running at. Read-only, and the standard way for runtime code to discover the privilege it is running under.
Floating-point control register: the knobs that change how arithmetic behaves rather than recording what happened. Rounding mode, denormal handling and exception trap enables all live here, and they apply to both scalar and Advanced SIMD instructions.
Floating-point status and control: the record of what has already happened. Each exception bit here is cumulative, meaning it records that the condition occurred at some point since it was last cleared, rather than the state of the most recent result.
Controls whether floating-point, Advanced SIMD and vector instructions are permitted to execute, and which levels may use them. EL1 software normally clears the FPEN field to its reset value and thus has the hardware trap every floating-point instruction to itself.
The system control register for EL1: one flag per subsystem, deciding whether the MMU runs, whether caches are on, which unprivileged system registers EL0 may touch, and how strict EL1 wants its own memory accesses to be. The register whose change has the largest immediate effect on subsequent instruction execution.
Root of the page table walk for the lower half of the address space, the region TTBR1_EL1 does not cover. The low bits of the register are configuration borrowed from TCR_EL1, so the two are read together.
Root of the page table walk for the upper half of the address space, conventionally the range that kernel code lives in. Separate from TTBR0_EL1 so a process can map its own pages low down while the kernel stays high up.
Describes the shape of the page tables that TTBR0_EL1 and TTBR1_EL1 point at: how big an address space each covers, how the walk caches memory, and whether tag checks apply. A translation regime only makes sense when this and both base registers agree.
Memory attribute indirection: sixteen attribute bytes that give names to the combinations of device, normal memory, cacheability and shareability that a page table entry can refer to. Each byte index becomes a value in the Attr field of a descriptor.
Where the exception vector table starts. On taking an exception to EL1 the processor loads the base of whichever of the 16 vector slots applies and branches there, so relocating all handlers means writing this one register.
The syndrome register: why the exception happened. A handler reads EC to learn the class of exception, then interprets ISS, whose layout depends entirely on that class. IL records whether the faulting instruction was 32 or 64 bits, which is the first thing to check when stepping through an unknown exception.
Fault address register: the address that caused a memory abort. Only meaningful for the exception classes that actually supply one, which the ESR EC field tells you.
Exception link register: the address to return to when the handler finishes. Written by the hardware on exception entry, and rewritten to point at the faulting instruction before an exception return so execution can resume.
Saved program state register: a snapshot of PSTATE taken on exception entry and restored by exception return. Together with ELR_EL1 it holds everything needed to resume the interrupted code.
The stack pointer in force while EL0 code runs. EL1 software switches between its own SP and this one by writing SPSEL.
The stack pointer belonging to EL1. Higher levels can read it but not write it, which is how a hypervisor inspects a guest kernel's stack.
The stack pointer belonging to EL2, the hypervisor.
The stack pointer belonging to EL3. Reachable only from EL3, and the only stack pointer register with no lower-privilege reader.
The main implementer ID register: who made the chip and what model it is. Read-only from every level, and the first thing a program checks to work out what it is running on.
Multiprocessor ID: which core in the cluster is running this code. How a scheduler works out which thread it is on, and how an interrupt handler finds the core it belongs to.
An operating-system-chosen identifier for the running process. A debugger or tracer reads it to attribute samples to the right thread; the architecture cares nothing about what the value means.
Software thread ID register, readable from EL0. A thread library keeps the per-thread pointer here so that a call can find the current thread without a system call.
The read-only companion to TPIDR_EL0. A runtime can leave it unwritable by user code so a signal handler can trust it while still reading the thread pointer.
Cache topology register: the parameters software needs in order to clean and invalidate caches correctly. Read from EL0, so an interrupt handler can maintain the cache without a system call.
The first of the architectural feature identification registers, each nibble answering one yes-or-no question about the processor. Software reads these at startup to decide which code paths to take.
Whether the generic timer counts, and whether EL0 can read it. Bare-metal code reads this on boot to learn the counter frequency, because the counters themselves do not say how fast they tick.
The physical counter: ticks at a fixed frequency and never stops, whatever the CPU is doing. Read twice around the work you want to time and subtract.
The virtual counter: the same ticks as CNTPCT_EL0, offset by CNTVOFF_EL2. A guest reading it sees its own clock rather than the machine's, which is what makes time in a virtualised guest stable across migration.
The offset between the virtual and physical counters. A hypervisor writes it when migrating a guest, so the guest's clock does not jump.
Which cache to ask about next. Write a level and a type, then read CLIDR_EL1 or the cache's own ID register. It is a selector rather than a register in its own right, and it only has an effect on the next CLSR.
Breakpoint control: what address to match on, whether it is an instruction or a data access, and which privilege levels it fires at. This is the register a debugger writes to plant a breakpoint; the paired DBGBVR holds the address.
A second breakpoint, with the same shape as DBGBCR0_EL1 and its own address in the paired value register. Enough breakpoints to cover a handful of locations, which is what a watchpoint-heavy debugging session needs.
The address a breakpoint compares against. Paired with DBGBCR0_EL1: the control register says how to match, this says what to match.
Where to resume at EL2 after an exception. Set it to the faulting address to retry, which is what every exception handler does on return, or to the next instruction to move on.
Where to resume at EL3 after an exception.
Why the exception happened at EL2. Same layout as ESR_EL1: EC names the class, and ISS means something different for each of them.
Why the exception happened at EL3. The first register a secure-monitor handler reads.
The address that faulted at EL2. Only meaningful when the ESR's FnV bit says the fault address is valid, which is the first thing to check when a handler reads ESR and gets nothing useful from FAR.
The address that faulted at EL3. Check the ESR's FnV bit before trusting it.
Which architectural feature extensions are present. A bit here means the CPU implements an optional feature, and reading this is how code decides whether it may use one.
Which debug features are present: how many breakpoints and watchpoints, whether the debug architecture is v8 or later, and whether secure-state debugging is available.
Further debug features, including the self-hosted debug extensions a debugger uses when it is not attached over JTAG.
Which instruction-set extensions are present: atomics, CRC, the cryptographic instructions, the memory-tagging instructions. A program that checks this can use an extension when the CPU has it and fall back when it does not.
Further instruction-set extensions, including the memory-ordering instructions that pair with RCpc.
Which memory-management features are present, and the most important field on the whole register: PARange, the physical address width. A program that writes this register's PARange field is how Linux discovers how wide the address bus is.
Further memory-management features, including the tagged-address and address-translation-service extensions.
Which A64 floating-point features are present, beyond the first register's set: bigger vectors, more vector lengths, half-precision support. A feature-detect routine reads this to find out what it is allowed to use.
What the memory attributes in the page tables mean: eight 8-byte entries, each an index into this table. A descriptor's AttrIndx picks one of them, so this is what decides whether a page is Normal cacheable or Device memory.
What the memory attributes mean at EL3. Eight 8-byte entries, one per attribute index.
How many breakpoint registers the implementation actually has. Read it before filling in a debugger's breakpoint table, because the architectural minimum is four and real parts often have more.
Lock access to the debug registers, in the same way ASLRL_EL1 locks the architectural feature registers: write a magic value to lock, and a different one to unlock. Once locked, a debugger cannot change breakpoints behind the operating system's back.
The performance monitor's master control: whether counting is on, how many counters exist, and whether they can be programmed from EL0. Read N and P to find out what the chip actually offers, rather than assuming the architectural minimum.
A revision number the implementer assigns. The architecture does not say what it means, so two parts from the same vendor can report different values for otherwise identical hardware.
Whether the MMU and caches are on at EL2, and how this level handles exceptions. A hypervisor turns this on as the last step of setting up a guest.
Whether the MMU and caches are on at EL3. The secure monitor normally runs with translation off, since it has to be able to reach any physical address whatever the guests have done.
Which exception bank to use at EL2 or EL3: 0 for SP_ELx, 1 for SP_Elx. The upper bank is not used for normal exception entry, which is a distinction that costs people an afternoon the first time they meet it.
The processor state to restore when returning from an exception at EL2: which stack pointer, which exception mask bits were set, and the mode to return to. The saved PC lives in the matching ELR, not here.
The processor state to restore when returning from an exception at EL3.
How the MMU at EL2 translates: how many address bits are significant, the granule size for each regime, and where each root table is expected to be aligned.
How the MMU at EL3 translates. The secure monitor usually leaves translation off at EL3, so this is read more than written.
Per-thread storage at EL1. A kernel puts its current task pointer here, and TLS code reads it as the base of the thread's stack.
Per-thread storage at EL2, for a hypervisor doing the same trick TPIDR_EL1 is used for in a kernel.
Per-thread storage at EL3.
The root of the page tables used for the EL0 regime at EL2. A hypervisor points this at a table for a guest it is running.
The root of the page tables used for the EL0 regime at EL3.
The root of the page tables used for the EL1 regime at EL2. High addresses need this one rather than TTBR0_EL2, which is the distinction that catches people out when a guest faults on a kernel address.
Where this level's exception vectors start in memory. The layout after the base is fixed: one vector per exception class, at 0x208, 0x280 and so on. The same 16 slots exist at every level, so EL2 and EL3 have their own base.
Where the secure monitor's exception vectors start. Identical in layout to the other levels' VBARs, and encoded identically too, which is worth knowing when reading a decoder: the register name is the only thing that tells them apart.