Module: PWN::AI::Agent::Metrics
- Defined in:
- lib/pwn/ai/agent/metrics.rb
Overview
PWN::AI::Agent::Metrics is the telemetry layer of the pwn-ai learning loop. Every tool dispatch performed by PWN::AI::Agent::Loop is recorded here (name, success, duration, last error) and persisted to ~/.pwn/metrics.json.
PromptBuilder re-injects a compact effectiveness summary into the system prompt on every turn, so the model gains awareness of which tools historically succeed vs. fail on THIS host and can adapt its tool selection accordingly. This is one half of the closed feedback loop that lets pwn-ai continuously make itself smarter (the other half is PWN::AI::Agent::Learning).
PER-ENGINE SEGMENTATION
A local Ollama model and a frontier model do NOT have the same per-tool success rate — blending them mis-advises the local model about itself. Every record now also increments an :engines sub-bucket; summary/to_context accept engine: to surface only that engine's telemetry so the TOOL EFFECTIVENESS block becomes a genuine per-engine learned policy.
Constant Summary collapse
- METRICS_FILE =
File.join(Dir.home, '.pwn', 'metrics.json')
- HALF_LIFE_DAYS =
14.0- CUSUM_K =
0.15- CUSUM_H =
0.6- WINDOW =
30- PRM_MIN_N =
P18/P2 — rolling mean step_reward advantage for a tool. Sample-efficiency gate: < PRM_MIN_N samples → 0 (no rank noise). Shrinkage: adv *= min(1, n/PRM_FULL_N) so sparse signal cannot dominate UCB. Zero-variance windows (all +1 or all -1 from a single session) damp to 0.5×.
5- PRM_FULL_N =
20
Class Method Summary collapse
-
.advantage(opts = {}) ⇒ Object
- Supported Method Parameters
a = PWN::AI::Agent::Metrics.advantage(name: 'shell').
-
.authors ⇒ Object
- Author(s)
0day Inc.
-
.calibration(opts = {}) ⇒ Object
- Supported Method Parameters
cal = PWN::AI::Agent::Metrics.calibration(engine: :ollama).
-
.changepoints(opts = {}) ⇒ Object
- Supported Method Parameters
cps = PWN::AI::Agent::Metrics.changepoints.
-
.effective_rate(opts = {}) ⇒ Object
Blended success rate: when proxy_distrust > 0 and judge samples exist, mix judge_rate into the handler-ok rate.
-
.help ⇒ Object
Display Usage for this Module.
-
.judge_confidence(opts = {}) ⇒ Object
P1 — mean judge confidence for a tool (nil when no samples).
-
.judge_rate(opts = {}) ⇒ Object
Mean judge score for a tool (nil when no ORM samples yet).
-
.load ⇒ Object
- Supported Method Parameters
metrics = PWN::AI::Agent::Metrics.load.
- .prm_advantage(opts = {}) ⇒ Object
- .prm_n(opts = {}) ⇒ Object
-
.proxy_trust ⇒ Object
P4 helper — Registry.rank calls this so β·advantage is scaled down when the proxy is untrustworthy.
-
.record(opts = {}) ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.record( name: 'required - tool name that was dispatched', success: 'required - Boolean, did the handler complete without error', duration: 'optional - Float seconds the dispatch took', error: 'optional - String error message when success is false', engine: 'optional - Symbol/String AI engine that chose this tool (segments telemetry)' ).
-
.record_calibration(opts = {}) ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.record_calibration(predicted:, actual:, brier:, engine:).
-
.record_judge(opts = {}) ⇒ Object
P20 — fold ORM judge (0..1) into per-tool telemetry so UCB / Thompson / advantage can prefer judge-grounded rates over the inflated handler-ok proxy when proxy_distrust is high.
-
.record_step_reward(opts = {}) ⇒ Object
P18 — called by Reward.prm after session annotate so live routing can bias toward tools that recently advanced goals (+1 step_reward).
-
.reset ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.reset.
-
.save(opts = {}) ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.save( metrics: 'required - Hash returned by .load / mutated in place' ).
-
.summary(opts = {}) ⇒ Object
- Supported Method Parameters
rows = PWN::AI::Agent::Metrics.summary( limit: 'optional - cap number of tools returned (default 25)', engine: 'optional - only that engine's sub-bucket (falls back to global when absent)' ).
-
.thompson(opts = {}) ⇒ Object
- Supported Method Parameters
p = PWN::AI::Agent::Metrics.thompson(name: 'shell').
-
.to_context(opts = {}) ⇒ Object
- Supported Method Parameters
ctx = PWN::AI::Agent::Metrics.to_context( limit: 'optional - cap number of tools included (default 8)', engine: 'optional - restrict to one engine's telemetry' ).
- .ucb(opts = {}) ⇒ Object
Class Method Details
.advantage(opts = {}) ⇒ Object
- Supported Method Parameters
a = PWN::AI::Agent::Metrics.advantage(name: 'shell')
C1 — tool.success_rate − global_rate over the rolling window.
243 244 245 246 247 248 249 250 251 252 253 254 255 256 |
# File 'lib/pwn/ai/agent/metrics.rb', line 243 public_class_method def self.advantage(opts = {}) data = load[:tools] || {} t = data[opts[:name].to_s.to_sym] return 0.0 unless t # P20 — local/global from effective_rate so bandit tracks judge # when the handler-ok proxy is hacked (proxy_distrust high). local = effective_rate(name: opts[:name]) rates = data.keys.map { |k| effective_rate(name: k) } global = rates.empty? ? 0.5 : (rates.sum / rates.length) (local - global).round(3) rescue StandardError 0.0 end |
.authors ⇒ Object
- Author(s)
0day Inc. support@0dayinc.com
535 536 537 |
# File 'lib/pwn/ai/agent/metrics.rb', line 535 public_class_method def self. "AUTHOR(S):\n 0day Inc. <support@0dayinc.com>\n" end |
.calibration(opts = {}) ⇒ Object
- Supported Method Parameters
cal = PWN::AI::Agent::Metrics.calibration(engine: :ollama)
446 447 448 449 450 451 452 453 |
# File 'lib/pwn/ai/agent/metrics.rb', line 446 public_class_method def self.calibration(opts = {}) eng = (opts[:engine] || :global).to_s.to_sym c = (load[:calibration] || {})[eng] return { n: 0, brier: nil } unless c && c[:n].to_i.positive? n = c[:n].to_f { n: c[:n], brier: (c[:brier_sum] / n).round(4), mean_predicted: (c[:p_sum] / n).round(3), mean_actual: (c[:a_sum] / n).round(3), overconfidence: ((c[:p_sum] - c[:a_sum]) / n).round(3) } end |
.changepoints(opts = {}) ⇒ Object
- Supported Method Parameters
cps = PWN::AI::Agent::Metrics.changepoints
E1 — tools whose CUSUM tripped (success_rate regime change). The caller (Mistakes.record / Curriculum) triggers extro_snapshot + correlate on these so a Mistake caused by env drift is tagged cause: :env_drift and does NOT count toward [REPEATING].
410 411 412 413 414 415 416 417 418 419 420 421 |
# File 'lib/pwn/ai/agent/metrics.rb', line 410 public_class_method def self.changepoints(opts = {}) within = (opts[:within_secs] || 3_600).to_i now = Time.now.utc (load[:tools] || {}).filter_map do |name, t| cp = t[:changepoint_at] next unless cp && (now - Time.parse(cp)) < within { name: name.to_s, at: cp, window_rate: Array(t[:window]).sum.to_f / [Array(t[:window]).length, 1].max } end rescue StandardError [] end |
.effective_rate(opts = {}) ⇒ Object
Blended success rate: when proxy_distrust > 0 and judge samples exist, mix judge_rate into the handler-ok rate. distrust=1 → pure judge (or 0.5 if no judge data). distrust=0 → pure proxy.
373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 |
# File 'lib/pwn/ai/agent/metrics.rb', line 373 public_class_method def self.effective_rate(opts = {}) name = opts[:name].to_s data = load[:tools] || {} t = data[name.to_sym] return 0.5 unless t calls = [t[:calls].to_f, 1.0].max proxy = t[:ok].to_f / calls win = Array(t[:window]) proxy = win.sum.to_f / win.length if win.length >= 3 distrust = 1.0 - proxy_trust jrate = judge_rate(name: name) if distrust > 0.05 && !jrate.nil? # P1 — scale judge weight by judge confidence so a thin local # heuristic ORM cannot fully replace proxy when distrust is high. # effective_distrust = distrust * mean(confidence), floor 0.15 when # we do have judge samples so the signal still moves the needle. jconf = judge_confidence(name: name) jconf = 0.7 if jconf.nil? eff_d = (distrust * jconf).clamp(0.0, 1.0) eff_d = [eff_d, 0.15].max if jconf >= 0.3 ((proxy * (1.0 - eff_d)) + (jrate * eff_d)).clamp(0.0, 1.0).round(3) else proxy.round(3) end rescue StandardError 0.5 end |
.help ⇒ Object
Display Usage for this Module
541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 |
# File 'lib/pwn/ai/agent/metrics.rb', line 541 public_class_method def self.help puts <<~USAGE USAGE: PWN::AI::Agent::Metrics.record(name: 'shell', success: true, duration: 0.42, engine: :ollama) PWN::AI::Agent::Metrics.summary(limit: 10, engine: :ollama) PWN::AI::Agent::Metrics.to_context(limit: 8, engine: :ollama) # injected by PromptBuilder PWN::AI::Agent::Metrics.ucb(name: 'shell') # C1 exploration bonus PWN::AI::Agent::Metrics.thompson(name: 'shell') # C1 Beta(ok+1,fail+1) sample PWN::AI::Agent::Metrics.advantage(name: 'shell') # C1 local − global (P20 judge-blended) PWN::AI::Agent::Metrics.record_judge(name: 'shell', score: 0.8) # P20 ORM→metrics PWN::AI::Agent::Metrics.judge_rate(name: 'shell') # P20 mean ORM PWN::AI::Agent::Metrics.effective_rate(name: 'shell') # P20 proxy⋈judge PWN::AI::Agent::Metrics.changepoints(within_secs: 3600) # E1 CUSUM regime changes PWN::AI::Agent::Metrics.record_calibration(predicted: 0.8, actual: 1.0, brier: 0.04, engine: :ollama) PWN::AI::Agent::Metrics.calibration(engine: :ollama) # W3 Brier / overconfidence PWN::AI::Agent::Metrics.reset PWN::AI::Agent::Metrics.load PWN::AI::Agent::Metrics.save(metrics: hash) #{self}.authors USAGE end |
.judge_confidence(opts = {}) ⇒ Object
P1 — mean judge confidence for a tool (nil when no samples).
345 346 347 348 349 350 351 352 353 354 355 |
# File 'lib/pwn/ai/agent/metrics.rb', line 345 public_class_method def self.judge_confidence(opts = {}) t = (load[:tools] || {})[opts[:name].to_s.to_sym] return nil unless t win = Array(t[:judge_conf_window]) return nil if win.empty? (win.sum.to_f / win.length).round(3) rescue StandardError nil end |
.judge_rate(opts = {}) ⇒ Object
Mean judge score for a tool (nil when no ORM samples yet).
358 359 360 361 362 363 364 365 366 367 368 |
# File 'lib/pwn/ai/agent/metrics.rb', line 358 public_class_method def self.judge_rate(opts = {}) t = (load[:tools] || {})[opts[:name].to_s.to_sym] return nil unless t win = Array(t[:judge_window]) return nil if win.empty? (win.sum.to_f / win.length).round(3) rescue StandardError nil end |
.load ⇒ Object
- Supported Method Parameters
metrics = PWN::AI::Agent::Metrics.load
40 41 42 43 44 45 46 47 |
# File 'lib/pwn/ai/agent/metrics.rb', line 40 public_class_method def self.load FileUtils.mkdir_p(File.dirname(METRICS_FILE)) return { tools: {}, updated_at: nil } unless File.exist?(METRICS_FILE) JSON.parse(File.read(METRICS_FILE), symbolize_names: true) rescue StandardError { tools: {}, updated_at: nil } end |
.prm_advantage(opts = {}) ⇒ Object
266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 |
# File 'lib/pwn/ai/agent/metrics.rb', line 266 public_class_method def self.prm_advantage(opts = {}) data = load[:tools] || {} t = data[opts[:name].to_s.to_sym] return 0.0 unless t win = Array(t[:prm_window]) n = win.length return 0.0 if n < PRM_MIN_N mean = win.sum.to_f / n globals = data.values.map { |v| Array(v[:prm_window]) }.select { |w| w.length >= PRM_MIN_N } gmean = if globals.empty? 0.0 else all = globals.flatten all.sum.to_f / all.length end adv = mean - gmean # shrinkage toward 0 until PRM_FULL_N shrink = [n.to_f / PRM_FULL_N, 1.0].min # variance damp: if all equal, halve influence uniq = win.uniq var_damp = uniq.length <= 1 ? 0.5 : 1.0 (adv * shrink * var_damp).round(3) rescue StandardError 0.0 end |
.prm_n(opts = {}) ⇒ Object
294 295 296 297 298 299 300 301 |
# File 'lib/pwn/ai/agent/metrics.rb', line 294 public_class_method def self.prm_n(opts = {}) t = (load[:tools] || {})[opts[:name].to_s.to_sym] return 0 unless t Array(t[:prm_window]).length rescue StandardError 0 end |
.proxy_trust ⇒ Object
P4 helper — Registry.rank calls this so β·advantage is scaled down when the proxy is untrustworthy.
190 191 192 193 194 195 |
# File 'lib/pwn/ai/agent/metrics.rb', line 190 public_class_method def self.proxy_trust d = defined?(Reward) && Reward.respond_to?(:proxy_distrust) ? Reward.proxy_distrust : 0.0 (1.0 - d.to_f).clamp(0.0, 1.0) rescue StandardError 1.0 end |
.record(opts = {}) ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.record( name: 'required - tool name that was dispatched', success: 'required - Boolean, did the handler complete without error', duration: 'optional - Float seconds the dispatch took', error: 'optional - String error message when success is false', engine: 'optional - Symbol/String AI engine that chose this tool (segments telemetry)' )
83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 |
# File 'lib/pwn/ai/agent/metrics.rb', line 83 public_class_method def self.record(opts = {}) name = opts[:name].to_s success = opts[:success] ? true : false duration = opts[:duration].to_f error = opts[:error] engine = opts[:engine].to_s return if name.empty? metrics = load metrics[:tools] ||= {} key = name.to_sym t = metrics[:tools][key] ||= blank_bucket bump(bucket: t, success: success, duration: duration, error: error) unless engine.empty? t[:engines] ||= {} e = t[:engines][engine.to_sym] ||= blank_bucket bump(bucket: e, success: success, duration: duration, error: error) end save(metrics: metrics) t end |
.record_calibration(opts = {}) ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.record_calibration(predicted:, actual:, brier:, engine:)
W3 — plan_first emits p(success); Loop.run calls this with the realised outcome. Tracked per-engine so calibration of the local LoRA vs frontier is comparable.
430 431 432 433 434 435 436 437 438 439 440 441 |
# File 'lib/pwn/ai/agent/metrics.rb', line 430 public_class_method def self.record_calibration(opts = {}) m = load m[:calibration] ||= {} eng = (opts[:engine] || :global).to_s.to_sym c = m[:calibration][eng] ||= { n: 0, brier_sum: 0.0, p_sum: 0.0, a_sum: 0.0 } c[:n] += 1 c[:brier_sum] += opts[:brier].to_f c[:p_sum] += opts[:predicted].to_f c[:a_sum] += opts[:actual].to_f save(metrics: m) c end |
.record_judge(opts = {}) ⇒ Object
P20 — fold ORM judge (0..1) into per-tool telemetry so UCB / Thompson / advantage can prefer judge-grounded rates over the inflated handler-ok proxy when proxy_distrust is high.
325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 |
# File 'lib/pwn/ai/agent/metrics.rb', line 325 public_class_method def self.record_judge(opts = {}) name = opts[:name].to_s return if name.empty? score = opts[:score].to_f.clamp(0.0, 1.0) conf = opts.key?(:confidence) ? opts[:confidence].to_f.clamp(0.0, 1.0) : 0.7 m = load m[:tools] ||= {} t = m[:tools][name.to_sym] ||= blank_bucket t[:judge_window] = (Array(t[:judge_window]) + [score]).last(40) t[:judge_conf_window] = (Array(t[:judge_conf_window]) + [conf]).last(40) t[:judge_sum] = t[:judge_window].sum.to_f t[:judge_n] = t[:judge_window].length save(metrics: m) t rescue StandardError nil end |
.record_step_reward(opts = {}) ⇒ Object
P18 — called by Reward.prm after session annotate so live routing can bias toward tools that recently advanced goals (+1 step_reward).
305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 |
# File 'lib/pwn/ai/agent/metrics.rb', line 305 public_class_method def self.record_step_reward(opts = {}) name = opts[:name].to_s return if name.empty? rew = opts[:reward].to_f.clamp(-1.0, 1.0) m = load m[:tools] ||= {} t = m[:tools][name.to_sym] ||= blank_bucket t[:prm_window] = (Array(t[:prm_window]) + [rew]).last(40) t[:prm_sum] = t[:prm_window].sum t[:prm_n] = t[:prm_window].length save(metrics: m) t rescue StandardError nil end |
.reset ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.reset
458 459 460 461 |
# File 'lib/pwn/ai/agent/metrics.rb', line 458 public_class_method def self.reset FileUtils.rm_f(METRICS_FILE) { tools: {}, updated_at: nil } end |
.save(opts = {}) ⇒ Object
- Supported Method Parameters
PWN::AI::Agent::Metrics.save( metrics: 'required - Hash returned by .load / mutated in place' )
54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 |
# File 'lib/pwn/ai/agent/metrics.rb', line 54 public_class_method def self.save(opts = {}) metrics = opts[:metrics] ||= { tools: {} } metrics[:updated_at] = Time.now.utc.iso8601 FileUtils.mkdir_p(File.dirname(METRICS_FILE)) # 4.4 — flock + atomic rename path = METRICS_FILE tmp = File.join(File.dirname(path), ".#{File.basename(path)}.#{Process.pid}.tmp") body = JSON.pretty_generate(metrics) File.open(tmp, File::WRONLY | File::CREAT | File::TRUNC, 0o644) do |f| f.flock(File::LOCK_EX) f.write(body) f.flush f.fsync end File.rename(tmp, path) metrics ensure FileUtils.rm_f(tmp) if defined?(tmp) && tmp && File.exist?(tmp) end |
.summary(opts = {}) ⇒ Object
- Supported Method Parameters
rows = PWN::AI::Agent::Metrics.summary( limit: 'optional - cap number of tools returned (default 25)', engine: 'optional - only that engine's sub-bucket (falls back to global when absent)' )
111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 |
# File 'lib/pwn/ai/agent/metrics.rb', line 111 public_class_method def self.summary(opts = {}) limit = opts[:limit] || 25 engine = opts[:engine].to_s tools = load[:tools] || {} rows = tools.map do |name, t| b = engine.empty? ? t : (t.dig(:engines, engine.to_sym) || t) calls = b[:calls].to_i ok = b[:ok].to_i rate = calls.positive? ? (ok.to_f / calls).round(3) : 0.0 avg = calls.positive? ? (b[:total_duration].to_f / calls).round(3) : 0.0 { name: name.to_s, calls: calls, success_rate: rate, judge_rate: judge_rate(name: name), effective_rate: effective_rate(name: name), avg_duration: avg, last_error: b[:last_error], last_at: b[:last_at] } end rows.reject { |r| r[:calls].zero? } .sort_by { |r| [-r[:calls], -r[:success_rate]] }.first(limit) end |
.thompson(opts = {}) ⇒ Object
- Supported Method Parameters
p = PWN::AI::Agent::Metrics.thompson(name: 'shell')
C1 — Thompson sample from Beta(ok+1, fail+1). Naturally balances exploit/explore; used by Registry.rank as the tie-breaker.
217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 |
# File 'lib/pwn/ai/agent/metrics.rb', line 217 public_class_method def self.thompson(opts = {}) t = (load[:tools] || {})[opts[:name].to_s.to_sym] || blank_bucket # P20 — when judge samples exist and distrust > 0, tilt Beta toward # judge_rate so Thompson explore/exploit tracks ORM not handler-ok. ok = t[:ok].to_f fail = t[:fail].to_f jn = Array(t[:judge_window]).length if jn >= 3 && proxy_trust < 0.95 jr = judge_rate(name: opts[:name]).to_f # pseudo-counts from judge window, mixed by distrust d = (1.0 - proxy_trust).clamp(0.0, 1.0) jok = jr * jn jfail = (1.0 - jr) * jn ok = (ok * (1.0 - d)) + (jok * d) fail = (fail * (1.0 - d)) + (jfail * d) end beta_sample(alpha: ok + 1.0, beta: fail + 1.0) rescue StandardError 0.5 end |
.to_context(opts = {}) ⇒ Object
- Supported Method Parameters
ctx = PWN::AI::Agent::Metrics.to_context( limit: 'optional - cap number of tools included (default 8)', engine: 'optional - restrict to one engine's telemetry' )
142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 |
# File 'lib/pwn/ai/agent/metrics.rb', line 142 public_class_method def self.to_context(opts = {}) limit = opts[:limit] || 8 engine = opts[:engine] rows = summary(limit: limit, engine: engine) return '' if rows.empty? # P4 — when Reward.sentinel says proxy is hacked, haircut displayed # success rates so the model does not trust the lie in the prompt. distrust = defined?(Reward) && Reward.respond_to?(:proxy_distrust) ? Reward.proxy_distrust : 0.0 scope = engine.to_s.empty? ? 'historical' : "engine=#{engine}" scope = "#{scope}, proxy_distrust=#{distrust.round(2)}" if distrust.positive? lines = rows.map do |r| # P20 — display effective_rate (judge-blended) when available; # fall back to distrust haircut on raw proxy. rate = r[:success_rate].to_f eff = r[:effective_rate] adj = if eff eff.to_f else rate - ((rate - 0.5) * distrust) end err = r[:last_error] ? " last_err=#{r[:last_error][0, 60]}" : '' jtag = r[:judge_rate] ? " judge=#{(r[:judge_rate].to_f * 100).round(1)}%" : '' tag = distrust.positive? || r[:judge_rate] ? ' (adj)' : '' " - #{r[:name]}: calls=#{r[:calls]} success=#{(adj * 100).round(1)}%#{tag}#{jtag} avg=#{r[:avg_duration]}s#{err}" end warn_line = distrust.positive? ? "WARNING: reward proxy diverges from judge — success rates haircut by distrust=#{distrust.round(2)}; prefer judge-scored exemplars over raw rates.\n" : '' # P0 ops — surface W1 generator_mix when diet is unhealthy so the # online controller (and the model) prefer underfilled sources. mix_line = '' if defined?(Reward) && Reward.respond_to?(:generator_mix) begin m = Reward.generator_mix unless m[:healthy] mix_line = "W1 MIX: n=#{m[:n]} traj=#{m[:trajectory_fraction]} " \ "urgent=#{Array(m[:urgent]).join(',')} " \ "suppress=#{Array(m[:suppress]).join(',')} " \ "rec=#{m[:recommendation]}\n" end rescue StandardError mix_line = '' end end "#{warn_line}#{mix_line}TOOL EFFECTIVENESS (#{scope}, adapt tool choice accordingly)\n#{lines.join("\n")}\n\n" end |
.ucb(opts = {}) ⇒ Object
197 198 199 200 201 202 203 204 205 206 207 208 209 |
# File 'lib/pwn/ai/agent/metrics.rb', line 197 public_class_method def self.ucb(opts = {}) name = opts[:name].to_s c = (opts[:c] || 1.4).to_f data = load[:tools] || {} t = data[name.to_sym] || blank_bucket n = [t[:calls].to_f, 1.0].max total = [data.values.sum { |v| v[:calls].to_f }, 1.0].max # P20 — mean from effective_rate (judge-blended when distrust high) mean = effective_rate(name: name) mean + (c * Math.sqrt(Math.log(total) / n)) rescue StandardError 1.0 end |