Module: Leakless::Cli::Triage
- Defined in:
- lib/leakless/triage.rb
Overview
Classifies findings as real leaks or false positives with an LLM.
SECURITY: the model never receives secret material, and there is no flag to opt out. Every payload line is rebuilt from the file on disk and passed through Redactor first, so what ships is the finding's shape: rule name, path, line number, length and entropy of the masked token, and whatever surrounding context survived redaction. Point --provider at Ollama to keep even that on the machine.
Verdicts are advisory. Triage annotates output; it never relaxes the high-severity gate, because a model must not be what decides a live key isn't a key.
Defined Under Namespace
Classes: Schema
Constant Summary collapse
- VERDICTS =
%w[true_positive false_positive uncertain].freeze
- OLLAMA_DEFAULT_BASE =
Ollama's conventional local endpoint, so --provider ollama needs no setup.
"http://localhost:11434/v1"- RATIONALE_LIMIT =
120- SOURCE_UNAVAILABLE =
"[source unavailable]"- FAILURES =
A missing API key or an unknown model id raise off a different branch of RubyLLM's hierarchy than provider/HTTP faults, and they are the two most likely failures here, so both branches are caught.
[RubyLLM::Error, RubyLLM::ConfigurationError, RubyLLM::ModelNotFoundError, Faraday::Error].freeze
- SYSTEM_PROMPT =
<<~PROMPT You triage secret-scanner findings for a tool that runs before someone publishes a directory. For each finding decide whether it is a real leaked credential (true_positive), a placeholder, example, test fixture or documentation sample (false_positive), or genuinely unclear (uncertain). Secret values have been redacted before reaching you. A masked token shows only its length and Shannon entropy, as "[REDACTED chars=20 entropy=3.8]". Judge from the rule name, the file path, the line number, and the surrounding context. High entropy and a config or source path suggest a real credential; low entropy, or paths like spec/, test/, fixtures/, docs/, README, *.example, *.sample suggest not. Never claim to have read the secret. Reply for every index you are given. PROMPT
Class Method Summary collapse
-
.call(findings, **chat_options) ⇒ Object
Returns a new array of findings, each with triage fields filled where the model answered.
-
.configure ⇒ Object
RubyLLM reads no provider credentials from the environment on its own, and a CLI has nowhere else to get them.
-
.entry(finding, index, lines) ⇒ Object
There is deliberately no fallback to
finding.previewwhen the file cannot be read: the preview is truncated, and a fragment cut mid-token matches neither its own rule nor the generic backstop, which is the exact leak reading from disk exists to prevent. - .parse(content, count) ⇒ Object
-
.prompt_for(findings) ⇒ Object
Rebuilds each finding's line from disk so redaction sees the whole line, never the already-truncated preview.
-
.rationale(raw) ⇒ Object
Models routinely blow past the length the schema asks for, and these land in a terminal one line per finding, so the limit is enforced here too.
- .read_lines(path) ⇒ Object
- .request(findings, **chat_options) ⇒ Object
-
.verdict_index(row, count) ⇒ Object
Only an in-range index paired with a label we asked for is usable; anything else is dropped so it cannot shift another finding's verdict.
Class Method Details
.call(findings, **chat_options) ⇒ Object
Returns a new array of findings, each with triage fields filled where the model answered. Findings the model skipped come back untouched.
71 72 73 74 75 76 77 78 79 |
# File 'lib/leakless/triage.rb', line 71 def call(findings, **) return findings if findings.empty? verdicts = request(findings, **) findings.each_with_index.map do |finding, index| verdict = verdicts[index] verdict ? finding.with(**verdict) : finding end end |
.configure ⇒ Object
RubyLLM reads no provider credentials from the environment on its own, and a CLI has nowhere else to get them.
95 96 97 98 99 100 101 |
# File 'lib/leakless/triage.rb', line 95 def configure RubyLLM.configure do |config| config.anthropic_api_key = ENV.fetch("ANTHROPIC_API_KEY", nil) config.openai_api_key = ENV.fetch("OPENAI_API_KEY", nil) config.ollama_api_base = ENV.fetch("OLLAMA_API_BASE", OLLAMA_DEFAULT_BASE) end end |
.entry(finding, index, lines) ⇒ Object
There is deliberately no fallback to finding.preview when the file cannot be
read: the preview is truncated, and a fragment cut mid-token matches neither its
own rule nor the generic backstop, which is the exact leak reading from disk
exists to prevent. Losing one finding's context is the cheaper failure. The path
is redacted too, so a token embedded in a filename cannot ride along.
116 117 118 119 120 121 122 123 |
# File 'lib/leakless/triage.rb', line 116 def entry(finding, index, lines) raw = lines[finding.file]&.[](finding.line - 1) <<~ENTRY #{index}. rule=#{finding.rule} severity=#{finding.severity} location=#{Redactor.redact(finding.file)}:#{finding.line} context=#{raw ? Redactor.redact(raw.chomp) : SOURCE_UNAVAILABLE} ENTRY end |
.parse(content, count) ⇒ Object
131 132 133 134 135 136 137 138 139 140 141 142 |
# File 'lib/leakless/triage.rb', line 131 def parse(content, count) return {} unless content.is_a?(Hash) Array(content["verdicts"]).each_with_object({}) do |row, acc| index = verdict_index(row, count) next unless index acc[index] = { verdict: row["verdict"], confidence: row["confidence"].to_f.clamp(0.0, 1.0), rationale: rationale(row["rationale"]) } end end |
.prompt_for(findings) ⇒ Object
Rebuilds each finding's line from disk so redaction sees the whole line, never the already-truncated preview.
105 106 107 108 109 |
# File 'lib/leakless/triage.rb', line 105 def prompt_for(findings) lines = Hash.new { |cache, path| cache[path] = read_lines(path) } findings.each_with_index.map { |finding, index| entry(finding, index, lines) }.join("\n") end |
.rationale(raw) ⇒ Object
Models routinely blow past the length the schema asks for, and these land in a terminal one line per finding, so the limit is enforced here too.
146 147 148 149 150 151 |
# File 'lib/leakless/triage.rb', line 146 def rationale(raw) text = raw.to_s.strip return text if text.length <= RATIONALE_LIMIT "#{text[0, RATIONALE_LIMIT - 1]}…" end |
.read_lines(path) ⇒ Object
125 126 127 128 129 |
# File 'lib/leakless/triage.rb', line 125 def read_lines(path) File.readlines(path, encoding: "UTF-8") rescue ArgumentError, SystemCallError nil end |
.request(findings, **chat_options) ⇒ Object
81 82 83 84 85 86 87 88 89 90 91 |
# File 'lib/leakless/triage.rb', line 81 def request(findings, **) configure chat = RubyLLM.chat(**.compact) .with_instructions(SYSTEM_PROMPT) .with_schema(Schema) # On 1.16 a schema-backed reply arrives parsed on #content (there is no # #parsed), and stays a String if the provider returned unparseable JSON. parse(chat.ask(prompt_for(findings)).content, findings.size) rescue *FAILURES => e raise Error, "AI triage failed: #{e.class}: #{e.}" end |
.verdict_index(row, count) ⇒ Object
Only an in-range index paired with a label we asked for is usable; anything else is dropped so it cannot shift another finding's verdict.
155 156 157 158 159 160 161 162 163 |
# File 'lib/leakless/triage.rb', line 155 def verdict_index(row, count) return unless row.is_a?(Hash) index = row["index"] return unless index.is_a?(Integer) && index.between?(0, count - 1) return unless VERDICTS.include?(row["verdict"]) index end |