資安週報Security Weekly 攻擊手法通報 × 資安工具Advisories × Tooling
AI 應用安全AI application security MIT

promptfoo

對 LLM 應用做紅隊測試與回歸測試的工具,把提示注入、資料外洩等風險變成可重複執行的測項。Red-teaming and regression testing for LLM applications, turning prompt injection and data-leak risks into repeatable test cases.

這是什麼What it is

LLM 應用的麻煩在於輸出不確定:同一個提示兩次執行結果可能不同,傳統的斷言式測試派不上用場。改了一句系統提示、換了模型版本,你很難知道有沒有弄壞什麼、或是打開了新的攻擊面。

promptfoo 用測試矩陣的方式處理這件事:定義一組輸入與期望性質,跨多個模型與多個提示版本批次執行,再用比對規則或模型評分來判定。內建的紅隊模組會自動產生對抗性輸入——提示注入、越獄、個資誘出、有害內容——並依 OWASP LLM Top 10 之類的分類整理結果。

它可以在本機跑,不必把資料送到第三方服務,也能接進 CI 當作上線前的關卡。

The awkward thing about LLM applications is that output is non-deterministic: the same prompt can return different results, so assertion-style tests do not transfer. Change one line of a system prompt or bump a model version and it is hard to know what broke — or what new attack surface opened.

promptfoo handles this as a test matrix: define inputs and expected properties, run them across multiple models and prompt versions, and judge with matchers or model-graded rubrics. Its red-team module generates adversarial inputs automatically — prompt injection, jailbreaks, PII extraction, harmful content — and organises findings against taxonomies like the OWASP LLM Top 10.

It runs locally without shipping your data to a third-party service, and slots into CI as a pre-release gate.

什麼時候用得上When to use it

- 改動提示或換模型之前後跑一次,確認行為沒有退步,這是最日常也最有價值的用法
- 上線前的紅隊掃描:讓它自動產生對抗性輸入,看你的防護提示與輸出過濾擋不擋得住
- 比較模型:同一組測項跨模型執行,用實際數據而非直覺決定要用哪一個
- CI 關卡:把通過率設為門檻,不達標就擋下部署

若你的應用有工具呼叫(agent、MCP)或會讀取外部內容,間接提示注入是最該優先測的項目——那是目前最難防也最常被忽略的一類。

- Run before and after prompt or model changes to confirm behaviour has not regressed — the most routine and most valuable use
- Pre-release red teaming: let it generate adversarial inputs and see whether your guard prompts and output filters hold
- Model comparison: run one suite across models and choose on data rather than intuition
- CI gating: set a pass-rate threshold and block deployment below it

If your application calls tools (agents, MCP) or reads external content, indirect prompt injection is the first thing to test — currently the hardest to defend and the most often overlooked.

注意事項Caveats

測試會實際呼叫模型,紅隊掃描的用量可能不小,先確認 API 成本與速率限制。另外,測項通過不代表安全——它驗證的是你想到要測的風險,想不到的仍然存在。把它當作回歸防護網,不是安全證明。

Tests make real model calls, and red-team sweeps can consume a lot of tokens — check API cost and rate limits first. Also, passing does not mean secure: it validates the risks you thought to test, not the ones you did not. Treat it as a regression net, not a proof of safety.

安裝Install

npx promptfoo@latest init
# 或 / or
brew install promptfoo

快速上手Quick start

# 建立設定並執行測試矩陣
npx promptfoo@latest init
npx promptfoo@latest eval

# 檢視結果
npx promptfoo@latest view

# 紅隊掃描:自動產生對抗性輸入
npx promptfoo@latest redteam init
npx promptfoo@latest redteam run

出現在Featured in