National Cyber Warfare Foundation (NCWF)

ptxNinja for reverse engineering CUDA PTX virtual instruction set code in Binary Ninja


0 user ratings
2026-09-29 21:32:56
milo
Red Team (CNA)
"ptxNinja

ptxNinja is a Binary Ninja architecture plugin that lets analysts explore and decompile PTX, the virtual instruction set behind CUDA GPU kernels, during authorized reverse engineering work.








Toolseekbytes/ptxNinja — Binary Ninja architecture plugin adding PTX support for CUDA kernel analysis
CategoryReverse engineering / binary analysis plugin (Rust, GPL-3.0)
Primary UseLoading, navigating, and decompiling PTX GPU kernel binaries inside Binary Ninja for authorized research and integration into automated analysis pipelines
Safe UseIntended for authorized malware research, defensive analysis, and academic study of GPU code on samples you own or are licensed to examine
Telemetry NotePure offline static analysis — it runs entirely inside Binary Ninja on local files and leaves no network or target-side footprint

ptxNinja is an architecture plugin for Binary Ninja written by Nicolò Altamura (seekbytes) that adds support for PTX, the intermediate virtual instruction set used by CUDA-based NVIDIA GPUs. In practice, it takes the text-based PTX representation that tools like nvdisasm and cuobjdump emit and turns it into something Binary Ninja can model: navigable functions, control flow graphs, and the decompiler view practitioners rely on. This matters because GPU code has historically been a blind spot in mainstream reverse engineering suites, which concentrate on x86, ARM, and other CPU architectures. As GPU-accelerated workloads — including machine learning models and increasingly GPU-resident malware techniques — become more common, the ability to statically examine kernel code stops being a niche curiosity and becomes a genuine defensive capability. With 62 stars and a GPL-3.0 license, this is a focused research tool rather than a polished commercial product, and it is honest about its limitations, which we will get to.


The core value proposition, per the README, is threefold. First, it lets you explore PTX binaries while integrating analyses already available in the wild, meaning the plugin slots into the existing Binary Ninja workflow rather than replacing it. Second, it gives you navigation through GPU kernels and functions, which sounds trivial until you have tried to read a raw PTX dump of a large kernel. Third — and this is the point that will resonate with operators — it makes PTX programmatically accessible through Binary Ninja's API, so existing automated tooling, headless scripts, and batch analysis pipelines can suddenly consume CUDA code without bespoke parsers. The repository's topics (binaryninja, cuda, decompilation, ptx) confirm the intent: this is about bringing PTX into the modern decompilation ecosystem. The implementation language is listed as Rust, which suggests the heavy lifting — likely the PTX grammar and parsing — happens in a compiled component rather than pure Python.


Installation is deliberately low-friction. The plugin can be installed directly through Binary Ninja's plugin manager, which is the recommended path for anyone who just wants to try it against the bundled examples. For a manual setup, the README walks through cloning the repository into Binary Ninja's plugin folder and installing the ptx-parser dependency via pip. A sensible touch is the suggestion to use a Python virtual environment (python -m venv ptxninja-env) to isolate dependencies, though the author notes that if you do so, you must manually point Binary Ninja's settings at the environment's site-packages path. That is a common friction point with Binary Ninja plugins and it is good the documentation addresses it up front rather than leaving users to debug mysterious import errors.


What really distinguishes this repository from a bare parser release is the examples/ directory. The author ships eleven toy PTX kernels covering a spread of real computational patterns: attention.ptx for an attention mechanism, relu.ptx, elu.ptx, gele.ptx, and softmax.ptx for neural network activation functions, gemm.ptx for general matrix multiplication, layernorm.ptx for layer normalization, histogram.ptx, matrix.ptx, multi_tensor.ptx, and reduce_sum.ptx for parallel reduction. This selection is not arbitrary — it maps almost exactly onto the vocabulary of modern deep learning inference, and the screenshot on the repo shows an attention mechanism reverse engineered in PTX. The implicit thesis is that analysts will increasingly encounter CUDA kernels implementing model logic (or hiding logic inside model logic), and these kernels provide a safe, self-contained corpus for learning to read PTX and for regression-testing the plugin itself.


Internally, the hard problem ptxNinja solves is grammatical. PTX is fundamentally a text format, and the output of nvdisasm/cuobjdump already carries substantial annotation embedded in the binaries. The author's engineering choice — building a precise grammar to test the parser against that text — is the classic approach, but text-based IRs are notoriously fiddly, and the README is refreshingly candid about where the parser breaks down. Anyone planning to rely on this tool for real work should read the limitations section carefully, because several of the caveats affect control flow integrity, which is the backbone of any decompiler output you would actually trust.


The documented limitations deserve individual attention. Labels are occasionally mismatched with their target instruction, which may produce incorrect control flow visualization in Binary Ninja — a serious caveat, since wrong branch targets can lead an analyst to misread a kernel's logic entirely. Global statements declared outside function scope are not supported, and memory space allocated outside functions is not mapped, so kernels that rely on module-level state will be incompletely modeled. Not all atomic operations (atom) are fully supported, which matters because atomics are pervasive in parallel kernels for synchronization and shared-memory updates. Finally, symbol names are not demangled; the cuobjdump/cufilt beautification step has not been applied yet, so expect mangled identifiers in the UI until that lands. None of these are disqualifying, but they collectively mean conclusions drawn from ptxNinja output should be cross-checked against the raw PTX text.


From an authorized-workflow perspective, the natural home for ptxNinja is a malware research or code-audit lab. Scenarios where it earns its keep include examining CUDA payloads embedded in suspicious binaries (GPU-resident execution is an established evasion area), auditing the kernels bundled with proprietary inference engines you are licensed to assess, and academic study of how model architectures compile down to virtual GPU instructions. Because Binary Ninja exposes its full API to plugins, ptxNinja also becomes a building block: a headless script could batch-extract kernel structure from a corpus of samples, or a larger triage pipeline could flag samples containing PTX sections for deeper review. That programmatic angle is arguably the most durable contribution here — the GUI is nice, but the automation surface is what turns a viewer into a capability.


Defensively, there is little to worry about and much to like. ptxNinja performs purely static, offline analysis on files you load locally; it makes no network connections and touches no external systems. For blue-team visibility, its use is entirely contained within the analyst's workstation and Binary Ninja instance, so there is no target-side telemetry associated with it. The main operational note is supply-chain hygiene: it is a third-party plugin pulling in the ptx-parser package, so environments with strict controls should pin versions and review the dependency before installing it into an analysis VM that handles sensitive samples.


Comparing it to the surrounding ecosystem, the alternatives for PTX inspection are mostly NVIDIA's own tooling — nvdisasm, cuobjdump, and the CUDA toolkit — which give you annotated text dumps but no decompiler view, no cross-referencing, and no scripting hooks comparable to Binary Ninja's API. That gap is precisely what ptxNinja fills, and the fact that it presents PTX through the same interface analysts already use for CPU code lowers the adoption cost to nearly zero for existing Binary Ninja users. The trade-off is maturity: NVIDIA's tools are battle-tested on every PTX variant the compiler emits, while this grammar-based parser has known blind spots around labels, globals, and atomics.


In sum, seekbytes/ptxNinja is a specialized but genuinely useful addition to the reverse engineer's toolkit. It is best understood as an early-stage bridge between the CUDA ecosystem and the Binary Ninja analysis platform, with a thoughtful example corpus aimed at machine learning kernels, documented limitations, and a clear automation story. For authorized professionals who occasionally need to answer "what does this GPU kernel actually do?", it converts an opaque text dump into navigable, scriptable analysis — with the caveat that its control flow reconstruction is still imperfect and should be verified against the underlying PTX source. Anyone working at the intersection of GPU computing and security should watch this repository as the demangling and atomic-operation support mature.



Official project repository for seekbytes/ptxNinja.

Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.






Source: OffensiveSec
Source Link: https://www.offsecblog.com/2026/09/ptxninja-for-reverse-engineering-cuda.html


Comments
new comment
Nobody has commented yet. Will you be the first?
 
Forum
Red Team (CNA)



Copyright 2012 through 2026 - National Cyber Warfare Foundation - All rights reserved worldwide.