A Case Study of Stack Pointer Behavior During C++ Function Calls on x86-64

This repository contains the complete case study paper titled "A Case Study of Stack Pointer Behavior During C++ Function Calls on x86-64" by Aditya Harishankar Prasad. Overview The study investigates how the x86-64 stack pointer (RSP) and base pointer (RBP) behave during C++ function execution under the System V AMD64 ABI on Linux. By tracking stack frame layout, local variable allocation, register argument spillover, and function returns across four compiler optimization levels (-O0, -O1, -O2, and -O3), the paper evaluates how optimizing compilers alter stack memory and register usage. Methodology & Toolchain Target Architecture & ABI: Linux x86-64, System V AMD64 ABI Primary Compiler: GCC 14.1.0 Secondary Compiler: Clang 18.1.0 Language Standard: C++17 Optimization Levels Evaluated: -O0, -O1, -O2, -O3 Experimental Workloads: Six controlled programs covering simple leaf calls, local variable allocation, argument overflow beyond the six-register threshold, nested multi-frame calls, recursion, and optimization transitions. Key Findings -O0 Baseline Frame Setup: Every function establishes an explicit frame anchor with pushq %rbp and movq %rsp, %rbp. All locals and arguments are spilled to memory slots. Function parameters beyond the six-register limit are passed by the caller on the stack and accessed by the callee via positive offsets from RBP. -O1 Frame-Pointer Omission: GCC enables -fomit-frame-pointer by default. RBP is reclaimed for general register allocation, variables remain in registers, and leaf functions utilize the 128-byte red zone without adjusting RSP. -O2 and -O3 Recursion Elimination: The compiler applies aggressive inlining and constant folding. In recursive routines such as factorial, GCC transforms the recursion into an iterative scalar loop, removing stack frame growth and recursive call instructions entirely. Clang applies loop vectorization to achieve the same result. ABI Requirements vs. Compiler Choices: The study separates mandatory ABI rules (16-byte stack alignment, register-based parameter passing via RDI-R9, RAX return values, and callee-saved register preservation) from compiler heuristics (frame-pointer omission, red zone usage, inlining thresholds, and tail-call replacement). Included Materials The deposit includes the complete academic paper in PDF format (4 pages, two-column layout), complete with: Four technical architecture diagrams (workflow, stack layout, call lifecycle, and optimization progression) Four evidence-based comparison tables Inspected assembly listings from real GCC and Clang compilations Traceable citations to official ABI specifications and systems literature

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23064319
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Case Study of Stack Pointer Behavior During C++ Function Calls on x86-64

aditya prasad
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
article

A Case Study of Stack Pointer Behavior During C++ Function Calls on x86-64

aditya prasad
article en

Abstract

This repository contains the complete case study paper titled "A Case Study of Stack Pointer Behavior During C++ Function Calls on x86-64" by Aditya Harishankar Prasad. Overview The study investigates how the x86-64 stack pointer (RSP) and base pointer (RBP) behave during C++ function execution under the System V AMD64 ABI on Linux. By tracking stack frame layout, local variable allocation, register argument spillover, and function returns across four compiler optimization levels (-O0, -O1, -O2, and -O3), the paper evaluates how optimizing compilers alter stack memory and register usage. Methodology & Toolchain Target Architecture & ABI: Linux x86-64, System V AMD64 ABI Primary Compiler: GCC 14.1.0 Secondary Compiler: Clang 18.1.0 Language Standard: C++17 Optimization Levels Evaluated: -O0, -O1, -O2, -O3 Experimental Workloads: Six controlled programs covering simple leaf calls, local variable allocation, argument overflow beyond the six-register threshold, nested multi-frame calls, recursion, and optimization transitions. Key Findings -O0 Baseline Frame Setup: Every function establishes an explicit frame anchor with pushq %rbp and movq %rsp, %rbp. All locals and arguments are spilled to memory slots. Function parameters beyond the six-register limit are passed by the caller on the stack and accessed by the callee via positive offsets from RBP. -O1 Frame-Pointer Omission: GCC enables -fomit-frame-pointer by default. RBP is reclaimed for general register allocation, variables remain in registers, and leaf functions utilize the 128-byte red zone without adjusting RSP. -O2 and -O3 Recursion Elimination: The compiler applies aggressive inlining and constant folding. In recursive routines such as factorial, GCC transforms the recursion into an iterative scalar loop, removing stack frame growth and recursive call instructions entirely. Clang applies loop vectorization to achieve the same result. ABI Requirements vs. Compiler Choices: The study separates mandatory ABI rules (16-byte stack alignment, register-based parameter passing via RDI-R9, RAX return values, and callee-saved register preservation) from compiler heuristics (frame-pointer omission, red zone usage, inlining thresholds, and tail-call replacement). Included Materials The deposit includes the complete academic paper in PDF format (4 pages, two-column layout), complete with: Four technical architecture diagrams (workflow, stack layout, call lifecycle, and optimization progression) Four evidence-based comparison tables Inspected assembly listings from real GCC and Clang compilations Traceable citations to official ABI specifications and systems literature

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 7%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.