National Cyber Warfare Foundation (NCWF)

Meta_SecAlign for built-in prompt injection defense in open-weight LLMs


0 user ratings
2026-10-02 01:27:56
milo
Red Team (CNA)
"Meta_SecAlign

facebookresearch/Meta_SecAlign releases open-weight LLMs with built-in prompt injection defense plus the full training and evaluation code, intended for defensive research and authorized security testing of agentic systems.








Toolfacebookresearch/Meta_SecAlign — open-weight foundation LLMs with built-in prompt injection defense, plus training and evaluation code
CategoryLLM security / defensive fine-tuning research framework (Python)
Primary UseDeploying and evaluating prompt-injection-resistant models like Meta-SecAlign-8B and Meta-SecAlign-70B on security benchmarks such as AgentDojo and InjecAgent
Safe UseDefensive research, red-team evaluation of models you own or are authorized to test, and building hardened agents in lab environments
Telemetry NotePurely defensive research artifact; evaluation runs require externally configured OpenAI/Gemini API keys in data/openai_configs.yaml and data/gemini_configs.yaml, which leave provider-side usage traces

Meta_SecAlign from facebookresearch is a defensive research project that answers one of the more persistent problems in modern agent security: prompt injection. Rather than bolting a filter or a guardrail onto an existing model, the project ships foundation LLMs — Meta-SecAlign-8B and Meta-SecAlign-70B — that were trained with a security-aware recipe the authors call SecAlign++, so resistance to injection is baked into the weights themselves. The README claims the 70B variant is the first fully open-source commercial-grade LLM with built-in prompt injection defense, comparable to gpt-5 and gemini-3-pro in agentic tool and web security, and that the training recipe incurs no noticeable drop across a broad set of utility benchmarks. For security professionals building or assessing LLM agents, this is a rare case where the defensive artifact, the training code, and the evaluation harness all live in one repository.


The architecture story is what makes the release interesting beyond the model weights. The training recipe is described as significantly improved from the earlier SecAlign project, also from facebookresearch, and produces a robust LoRA adapter on top of base models meta-llama/Llama-3.1-8B-Instruct and meta-llama/Llama-3.3-70B-Instruct. A particularly clever detail is the lora_alpha parameter exposed at test time: the default value of 8 uses the model exactly as trained, while values between 0 and 8 interpolate between the undefended base model and the defended one, giving operators a tunable utility–security dial rather than a binary choice. Extrapolating beyond 8 is possible but explicitly untested by the authors, which is a candid admission worth respecting before experimenting in production.


Getting the environment running is deliberately reproducible. The README specifies hardware requirements honestly: Meta-SecAlign-8B needs 4×80 GB A100s for training and a single 16 GB GPU for evaluation, while the 70B model demands 8×141 GB H200s to train and roughly four to eight 80 GB A100s to evaluate. The tooling is built around uv, the fast Python package manager, with a Python 3.13 virtual environment, requirements.txt for core dependencies, and a pinned torchtune==0.6.0 from the CUDA 12.6 index. A setup.py step then installs the data dependencies, including benchmark material, and downloads the two defended checkpoints from Hugging Face.


The evaluation harness is arguably the most reusable component for practitioners who never intend to fine-tune anything. run_tests.py orchestrates the full reproduction pipeline by sequentially invoking tests.py, test_lm_eval.py, test_agentdojo.py, and test_injecagent.py, logging everything into [model_path]/summary.tsv. The harness supports local models served through vLLM, a range of Hugging Face open-weight checkpoints, OpenAI models including gpt-4o-mini, gpt-4o, and gpt-5 (with a --gpt5_reasoning_effort flag defaulting to high), and Google models from gemini-2.0-flash through gemini-3-pro-preview. That breadth means the repository doubles as a standardized prompt-injection benchmarking rig you can point at nearly any modern model.


On the benchmark side the coverage is unusually comprehensive, which matters because injection defenses are notorious for trading utility for safety. Six security suites are supported: instruction-following tests via AlpacaFarm-Hacked, SEP, TaskTracker, and CyberSecEval2, plus agentic tool-calling suites InjecAgent and AgentDojo. These are balanced against eight utility benchmarks drawn from lm-evaluation-harness — MMLU, MMLU-Pro, BBH, IFEval, and GPQA Diamond — alongside AlpacaEval2, SEP utility scoring, and the AgentDojo utility track. The README notes that in SEP utility evaluation, AlpecaEval2 prompting is used against reference responses from meta-llama/Meta-Llama-3-8B-Instruct, a detail that signals care about apples-to-apples comparison methodology.


For teams that want to apply the defense to their own deployments, secalign_plus_plus.py provides the fine-tuning entry point. Out of the box it defensively fine-tunes meta-llama/Llama-3.1-8B-Instruct, with the 70B variant available by uncommenting the relevant line. A minimal demo.py lets you feed the models your own samples, your own injection attempts, or even run them against your codebase as a quick sanity check before committing to a full evaluation run. Both entry points are straightforward invocations: python demo.py for exploration and python run_tests.py -m [model_path] --lora_alpha [lora_alpha] for the full suite.


Licensing is nuanced and worth reading carefully before any commercial adoption. The README states the models themselves are licensed for commercial use under the Llama community licenses, while the codebase is non-commercial — the majority under CC-BY-NC — with embedded portions from AgentDojo, TaskTracker, and lm-evaluation-harness under MIT. The repository metadata flags the license as NOASSERTION, which reflects that mixed provenance. Organizations intending to ship products on top of this work should separate the model weights (commercially usable) from the training and evaluation code (research-only) in their compliance review, and the software is noted as deposited in the BAIR Open Research Commons in 2025.


In an authorized workflow, this project fits naturally into both defensive engineering and offensive assessment contexts. Builders of agentic systems that consume untrusted web content or tool outputs can evaluate whether swapping in a Meta-SecAlign checkpoint raises the bar against injection before attackers get to embed instructions in retrieved documents. Red teams operating under proper authorization can use the same benchmark suite to quantify how a client's chosen model stack behaves under AgentDojo-style tool-calling attacks, producing evidence rather than anecdotes. Because the harness evaluates commercial models like gpt-5 and gemini-3-pro-preview alongside open ones, it also serves as a neutral comparison framework when procurement decisions hinge on injection resistance.


There are practical caveats to keep in mind. Evaluation against OpenAI and Gemini models requires configuring API keys in data/openai_configs.yaml and optionally data/gemini_configs.yaml, with an AzureOpenAI example provided, so cloud-based comparisons carry cost and data-egress considerations. The hardware floor for the 70B model is steep, meaning most teams will interact with Meta-SecAlign-8B or the hosted evaluation targets rather than retraining. And as with any security claim, the no-noticeable-utility-drop result is relative to the enumerated benchmarks; defenders should treat it as a strong baseline to verify against their own workloads. With modest traction at 71 stars and a paper backing the work, Meta_SecAlign is a credible, well-documented artifact in a defensive space that badly needs open alternatives to proprietary guardrails.



Official project repository for facebookresearch/Meta_SecAlign.

Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.






Source: OffensiveSec
Source Link: https://www.offsecblog.com/2026/10/metasecalign-for-built-in-prompt.html


Comments
new comment
Nobody has commented yet. Will you be the first?
 
Forum
Red Team (CNA)



Copyright 2012 through 2026 - National Cyber Warfare Foundation - All rights reserved worldwide.