National Cyber Warfare Foundation (NCWF)

Inside Hexestra, the AI-native IDE that puts operators and agents in one pentest workspace


0 user ratings
2026-09-27 09:25:57
milo
Red Team (CNA)
"Inside

Hexestra is an Electron-based penetration testing IDE where a human operator and an AI agent share the same browser, terminals, traffic capture, asset graph, tasks, and evidence for authorized engagements.








ToolABSllk/Hexestra — AI-native penetration testing IDE where human operators and AI agents share one project workspace
CategorySecurity testing IDE / operator-AI orchestration desktop application (TypeScript, Electron)
Primary UseRunning authorized penetration tests in a single workspace with shared browser, terminals, traffic capture, NetMap asset graph, and structured Evidence→Finding→Vulnerability→Report records
Safe UseIntended exclusively for authorized security testing — engagements with explicit permission, lab environments such as the documented Northstar Demo Lab, and scope-controlled assessments; the README warns against use on systems you do not own or are not authorized to assess
Telemetry NoteAll agent commands are audited, evidence records preserve raw output, and engagement state (scope, tasks, permissions, conversation branches) is persisted durably in the project folder — providing a complete accountability trail defenders and reviewers can inspect

Hexestra attacks a problem every operator who has tried to fold an LLM into a pentest workflow knows intimately: the AI lives in one terminal, your browser is over there, your proxy is somewhere else, and your evidence is scattered across notes files. The project, published at ABSllk/Hexestra under Apache-2.0 and written in TypeScript on Electron, collapses those fragments into a single project workspace where the human and the AI agent genuinely share the same operational surface. That means one browser, one set of terminal sessions, one traffic capture stream, one asset graph, and one evidence chain — not a chat window bolted onto a scanner.


The architecture decision that stands out first is that Hexestra does not implement its own agent. It orchestrates external CLIs — Claude Code or Codex — installed separately in a Native or WSL runtime, launching their App Server and reusing their authentication. This is a pragmatic choice: the agent layer benefits from upstream model improvements without the IDE chasing them, and the operator keeps whatever provider configuration they already have, including third-party Anthropic-compatible endpoints configured via ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN environment variables. The README's DeepSeek integration example makes clear this is provider-agnostic plumbing, not a lock-in.


Control is the real design thesis. Hexestra separates a permission mode — ASK, AUTO, or BYPASS — from a distinct autonomy level, while Rules of Engagement and technical safety boundaries persist independently of both. Critically, the README is honest that Scope labels are advisory: they guide agent prioritization but do not block commands or traffic. The operator can inspect, guide, approve, interrupt, or fully take over at any point, including taking over live shell sessions the agent is using. That human-takeover capability, plus per-command agent auditing, is what elevates this from an AI wrapper to something a professional could plausibly run on a real engagement.


The traffic layer is built on mitmproxy, with a mitmdump runtime bundled in packaged builds. The built-in browser routes through capture with inspection, interception, a Repeater-style replay capability, and evidence capture. Rather than replacing Burp Suite — a losing proposition — Hexestra offers an optional Burp Bridge: an authenticated loopback connection that mirrors completed HTTP exchanges into Burp's Target > Site map and, where supported, Organizer. The README notes a Burp extension API limitation preventing synthetic entries in live Proxy > HTTP history, which is the kind of precise technical disclosure that suggests the authors actually tested the integration rather than marketing it.


What will interest methodology-focused testers most is the structured progression model: raw output becomes Evidence, which links to a Finding, which validates into a Vulnerability record carrying severity, lifecycle, impact, and remediation, and finally rolls into a Report. Alongside this sits NetMap, a graph of typed assets with relationships, provenance, and active objectives that drives what the README calls graph-guided testing. The screenshots use a fictional Northstar Demo Lab with reserved example.test domains and synthetic data — a detail worth applauding, since it demonstrates the documentation itself models responsible targeting.


Session handling is broader than most IDEs attempt. Hexestra manages local, WSL, SSH, jump-host, and raw reverse-shell sessions, all shared between human and agent with takeover support. Conversation handling is similarly deliberate: branches are non-destructive, preserving both the original reasoning path and a canonical project state. Durable engagement state means reopening a project folder restores scope, tasks, NetMap, evidence, findings, reports, workspace tabs, permissions, and conversation branches — a meaningful operational continuity feature for multi-day engagements where context loss is a real cost.


The optional per-project Mihomo multi-hop egress integration deserves its own scrutiny, because it is where the engineering is most defensive. Proxy enforcement is isolated to the active project with no TUN or system proxy; terminal sessions receive HTTP_PROXY, HTTPS_PROXY, ALL_PROXY, NO_PROXY, and WSLENV — and the README explicitly states that programs ignoring these variables and opening raw sockets can bypass the terminal boundary. Claude API traffic and secondary egress from remote shell commands are outside this version's enforcement. Fail-closed behavior is the default: a missing node, invalid chain, or crashed runtime blocks managed egress rather than silently falling back to direct. A two-hop acceptance smoke is runnable via HEXESTRA_MIHOMO_PATH=/path/to/mihomo npm run test:proxy-smoke. This level of boundary honesty is rare and exactly what an operator needs to make an informed trust decision.


Setup is straightforward for an Electron project: Node.js 24 and npm, then npm ci and npm run electron:dev from the repository root. Source runs need mitmproxy installed separately (uv tool install mitmproxy), and the optional Burp Bridge builds with JDK 17 via npm run build:burp-bridge, producing resources/burp-bridge/hexestra-burp-bridge.jar loaded as a Java extension and paired through a loopback port and token. Verification tooling includes npm run audit:public, npm run check, and npm run electron:build, indicating a project that treats its own build hygiene as part of the product.


From a defensive-research perspective, Hexestra is as interesting for what it records as what it does. Every agent command is audited, evidence preserves raw HTTP responses, and vulnerability records carry full lifecycle context — meaning the tool generates the accountability trail that authorized testing is supposed to produce. A blue team reviewing an engagement, or an organization evaluating whether agent-assisted testing is acceptable, can inspect exactly what the agent was permitted to do and what it actually did. That is the difference between AI-assisted testing done professionally and done recklessly.


The responsible-use posture is explicit and repeated: the README's warning block restricts use to authorized testing, and the closing section states that destructive or privacy-impacting actions require appropriate approval, exported evidence should be treated as sensitive, and Hexestra does not replace professional judgment or accountability. At 43 stars, this is an early project — the sort of thing worth evaluating in a lab before trusting on client work — but the scope labeling, fail-closed proxy design, and durable audit state suggest authors who understand the operational gravity of putting an agent on a shared keyboard.


For operators already running Claude Code or Codex next to Burp Suite, PowerShell, WSL, and SSH, Hexestra is best understood as an orchestration layer that makes those pieces legible to each other and to a single human supervisor. It does not invent new attack capability; it restructures the engagement surface so that authorization, evidence, and agent control are first-class objects rather than afterthoughts. In a moment when agent-driven pentesting tools mostly optimize for autonomy, a tool that optimizes for supervisability is a noteworthy bet — and one that maps cleanly onto how authorized assessments are actually governed.



Official project repository for ABSllk/Hexestra.

Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.






Source: OffensiveSec
Source Link: https://www.offsecblog.com/2026/09/inside-hexestra-ai-native-ide-that-puts.html


Comments
new comment
Nobody has commented yet. Will you be the first?
 
Forum
Red Team (CNA)



Copyright 2012 through 2026 - National Cyber Warfare Foundation - All rights reserved worldwide.