Module: PWN::AI::Agent::Metrics

Defined in:
lib/pwn/ai/agent/metrics.rb

Overview

PWN::AI::Agent::Metrics is the telemetry layer of the pwn-ai learning loop. Every tool dispatch performed by PWN::AI::Agent::Loop is recorded here (name, success, duration, last error) and persisted to ~/.pwn/metrics.json.

PromptBuilder re-injects a compact effectiveness summary into the system prompt on every turn, so the model gains awareness of which tools historically succeed vs. fail on THIS host and can adapt its tool selection accordingly. This is one half of the closed feedback loop that lets pwn-ai continuously make itself smarter (the other half is PWN::AI::Agent::Learning).

PER-ENGINE SEGMENTATION

A local Ollama model and a frontier model do NOT have the same per-tool success rate — blending them mis-advises the local model about itself. Every record now also increments an :engines sub-bucket; summary/to_context accept engine: to surface only that engine's telemetry so the TOOL EFFECTIVENESS block becomes a genuine per-engine learned policy.

Constant Summary collapse

METRICS_FILE =
File.join(Dir.home, '.pwn', 'metrics.json')
HALF_LIFE_DAYS =
14.0
CUSUM_K =
0.15
CUSUM_H =
0.6
WINDOW =
30
PRM_MIN_N =

P18/P2 — rolling mean step_reward advantage for a tool. Sample-efficiency gate: < PRM_MIN_N samples → 0 (no rank noise). Shrinkage: adv *= min(1, n/PRM_FULL_N) so sparse signal cannot dominate UCB. Zero-variance windows (all +1 or all -1 from a single session) damp to 0.5×.

5
PRM_FULL_N =
20

Class Method Summary collapse

Class Method Details

.advantage(opts = {}) ⇒ Object

Supported Method Parameters

a = PWN::AI::Agent::Metrics.advantage(name: 'shell')

C1 — tool.success_rate − global_rate over the rolling window.



244
245
246
247
248
249
250
251
252
253
254
255
256
257
# File 'lib/pwn/ai/agent/metrics.rb', line 244

public_class_method def self.advantage(opts = {})
  data = load[:tools] || {}
  t    = data[opts[:name].to_s.to_sym]
  return 0.0 unless t

  # P20 — local/global from effective_rate so bandit tracks judge
  # when the handler-ok proxy is hacked (proxy_distrust high).
  local = effective_rate(name: opts[:name])
  rates = data.keys.map { |k| effective_rate(name: k) }
  global = rates.empty? ? 0.5 : (rates.sum / rates.length)
  (local - global).round(3)
rescue StandardError
  0.0
end

.authorsObject

Author(s)

0day Inc. support@0dayinc.com



583
584
585
# File 'lib/pwn/ai/agent/metrics.rb', line 583

public_class_method def self.authors
  "AUTHOR(S):\n  0day Inc. <support@0dayinc.com>\n"
end

.calibration(opts = {}) ⇒ Object

Supported Method Parameters

cal = PWN::AI::Agent::Metrics.calibration(engine: :ollama)



463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
# File 'lib/pwn/ai/agent/metrics.rb', line 463

public_class_method def self.calibration(opts = {})
  store = load[:calibration] || {}
  buckets = if opts.key?(:engine) && !opts[:engine].to_s.empty?
              [store[opts[:engine].to_s.to_sym]].compact
            else
              store.values
            end
  c = buckets.each_with_object({ n: 0, brier_sum: 0.0, p_sum: 0.0, a_sum: 0.0 }) do |b, acc|
    next unless b.is_a?(Hash)

    acc[:n] += b[:n].to_i
    acc[:brier_sum] += b[:brier_sum].to_f
    acc[:p_sum] += b[:p_sum].to_f
    acc[:a_sum] += b[:a_sum].to_f
  end
  return { n: 0, brier: nil } unless c[:n].positive?

  n = c[:n].to_f
  { n: c[:n], brier: (c[:brier_sum] / n).round(4), mean_predicted: (c[:p_sum] / n).round(3), mean_actual: (c[:a_sum] / n).round(3), overconfidence: ((c[:p_sum] - c[:a_sum]) / n).round(3) }
end

.changepoints(opts = {}) ⇒ Object

Supported Method Parameters

cps = PWN::AI::Agent::Metrics.changepoints

E1 — tools whose CUSUM tripped (success_rate regime change). The caller (Mistakes.record / Curriculum) triggers extro_snapshot + correlate on these so a Mistake caused by env drift is tagged cause: :env_drift and does NOT count toward [REPEATING].



427
428
429
430
431
432
433
434
435
436
437
438
# File 'lib/pwn/ai/agent/metrics.rb', line 427

public_class_method def self.changepoints(opts = {})
  within = (opts[:within_secs] || 3_600).to_i
  now = Time.now.utc
  (load[:tools] || {}).filter_map do |name, t|
    cp = t[:changepoint_at]
    next unless cp && (now - Time.parse(cp)) < within

    { name: name.to_s, at: cp, window_rate: Array(t[:window]).sum.to_f / [Array(t[:window]).length, 1].max }
  end
rescue StandardError
  []
end

.effective_rate(opts = {}) ⇒ Object

Blended success rate: when proxy_distrust > 0 and judge samples exist, mix judge_rate into the handler-ok rate. distrust=1 → pure judge (or 0.5 if no judge data). distrust=0 → pure proxy.



390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
# File 'lib/pwn/ai/agent/metrics.rb', line 390

public_class_method def self.effective_rate(opts = {})
  name = opts[:name].to_s
  data = load[:tools] || {}
  t = data[name.to_sym]
  return 0.5 unless t

  calls = [t[:calls].to_f, 1.0].max
  proxy = t[:ok].to_f / calls
  win = Array(t[:window])
  proxy = win.sum.to_f / win.length if win.length >= 3
  distrust = 1.0 - proxy_trust
  jrate = judge_rate(name: name)
  if distrust > 0.05 && !jrate.nil?
    # P1 — scale judge weight by judge confidence so a thin local
    # heuristic ORM cannot fully replace proxy when distrust is high.
    # effective_distrust = distrust * mean(confidence), floor 0.15 when
    # we do have judge samples so the signal still moves the needle.
    jconf = judge_confidence(name: name)
    jconf = 0.7 if jconf.nil?
    eff_d = (distrust * jconf).clamp(0.0, 1.0)
    eff_d = [eff_d, 0.15].max if jconf >= 0.3
    ((proxy * (1.0 - eff_d)) + (jrate * eff_d)).clamp(0.0, 1.0).round(3)
  else
    proxy.round(3)
  end
rescue StandardError
  0.5
end

.health_lineObject

Supported Method Parameters

PWN::AI::Agent::Metrics.reset



487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
# File 'lib/pwn/ai/agent/metrics.rb', line 487

public_class_method def self.health_line
  gap = nil
  if defined?(Reward) && Reward.respond_to?(:sentinel)
    s = Reward.sentinel
    gap = s[:gap_proxy_judge] if s.is_a?(Hash)
  end
  trend = if defined?(Curriculum) && Curriculum.respond_to?(:repeating_trend)
            Curriculum.repeating_trend
          else
            {}
          end
  traj = (Reward.generator_mix[:trajectory_fraction] if defined?(Reward) && Reward.respond_to?(:generator_mix))
  parked = (Mistakes.operator_inbox(limit: 50)[:count] if defined?(Mistakes) && Mistakes.respond_to?(:operator_inbox))
  "HEALTH judge-proxy-gap=#{gap || '-'} repeating=#{trend[:status] || '-'} " \
    "w1-traj=#{traj || '-'} parked-needs-human=#{parked || '-'}\n"
rescue StandardError
  ''
end

.helpObject

Display Usage for this Module



589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
# File 'lib/pwn/ai/agent/metrics.rb', line 589

public_class_method def self.help
  puts <<~USAGE
    USAGE:
      PWN::AI::Agent::Metrics.record(name: 'shell', success: true, duration: 0.42, engine: :ollama)
      PWN::AI::Agent::Metrics.summary(limit: 10, engine: :ollama)
      PWN::AI::Agent::Metrics.to_context(limit: 8, engine: :ollama)   # injected by PromptBuilder
      PWN::AI::Agent::Metrics.ucb(name: 'shell')                 # C1 exploration bonus
      PWN::AI::Agent::Metrics.thompson(name: 'shell')            # C1 Beta(ok+1,fail+1) sample
      PWN::AI::Agent::Metrics.advantage(name: 'shell')           # C1 local − global (P20 judge-blended)
      PWN::AI::Agent::Metrics.record_judge(name: 'shell', score: 0.8) # P20 ORM→metrics
      PWN::AI::Agent::Metrics.judge_rate(name: 'shell')         # P20 mean ORM
      PWN::AI::Agent::Metrics.effective_rate(name: 'shell')     # P20 proxy⋈judge
      PWN::AI::Agent::Metrics.changepoints(within_secs: 3600)    # E1 CUSUM regime changes
      PWN::AI::Agent::Metrics.record_calibration(predicted: 0.8, actual: 1.0, brier: 0.04, engine: :ollama)
      PWN::AI::Agent::Metrics.calibration(engine: :ollama)       # W3 Brier / overconfidence
      PWN::AI::Agent::Metrics.reset
      PWN::AI::Agent::Metrics.load
      PWN::AI::Agent::Metrics.save(metrics: hash)

      #{self}.authors
  USAGE
end

.judge_confidence(opts = {}) ⇒ Object

P1 — mean judge confidence for a tool (nil when no samples).



349
350
351
352
353
354
355
356
357
358
359
# File 'lib/pwn/ai/agent/metrics.rb', line 349

public_class_method def self.judge_confidence(opts = {})
  t = (load[:tools] || {})[opts[:name].to_s.to_sym]
  return nil unless t

  win = Array(t[:judge_conf_window])
  return nil if win.empty?

  (win.sum.to_f / win.length).round(3)
rescue StandardError
  nil
end

.judge_rate(opts = {}) ⇒ Object

Mean judge score for a tool (nil when no ORM samples yet). When per-sample sources exist, LLM ORM outweighs heuristic overlap so the proxy haircut tracks the outcome model.



364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
# File 'lib/pwn/ai/agent/metrics.rb', line 364

public_class_method def self.judge_rate(opts = {})
  t = (load[:tools] || {})[opts[:name].to_s.to_sym]
  return nil unless t

  win = Array(t[:judge_window])
  return nil if win.empty?

  src = Array(t[:judge_src_window])
  if src.length == win.length && defined?(Reward) && Reward.respond_to?(:judge_sample_weight)
    num = 0.0
    den = 0.0
    win.each_with_index do |score, i|
      wt = Reward.judge_sample_weight(source: src[i]).to_f
      num += score.to_f * wt
      den += wt
    end
    return (num / den).round(3) if den.positive?
  end
  (win.sum.to_f / win.length).round(3)
rescue StandardError
  nil
end

.loadObject

Supported Method Parameters

metrics = PWN::AI::Agent::Metrics.load



40
41
42
43
44
45
46
47
# File 'lib/pwn/ai/agent/metrics.rb', line 40

public_class_method def self.load
  FileUtils.mkdir_p(File.dirname(METRICS_FILE))
  return { tools: {}, updated_at: nil } unless File.exist?(METRICS_FILE)

  JSON.parse(File.read(METRICS_FILE), symbolize_names: true)
rescue StandardError
  { tools: {}, updated_at: nil }
end

.prm_advantage(opts = {}) ⇒ Object



267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
# File 'lib/pwn/ai/agent/metrics.rb', line 267

public_class_method def self.prm_advantage(opts = {})
  data = load[:tools] || {}
  t = data[opts[:name].to_s.to_sym]
  return 0.0 unless t

  win = Array(t[:prm_window])
  n = win.length
  return 0.0 if n < PRM_MIN_N

  mean = win.sum.to_f / n
  globals = data.values.map { |v| Array(v[:prm_window]) }.select { |w| w.length >= PRM_MIN_N }
  gmean = if globals.empty?
            0.0
          else
            all = globals.flatten
            all.sum.to_f / all.length
          end
  adv = mean - gmean
  # shrinkage toward 0 until PRM_FULL_N
  shrink = [n.to_f / PRM_FULL_N, 1.0].min
  # variance damp: if all equal, halve influence
  uniq = win.uniq
  var_damp = uniq.length <= 1 ? 0.5 : 1.0
  (adv * shrink * var_damp).round(3)
rescue StandardError
  0.0
end

.prm_n(opts = {}) ⇒ Object



295
296
297
298
299
300
301
302
# File 'lib/pwn/ai/agent/metrics.rb', line 295

public_class_method def self.prm_n(opts = {})
  t = (load[:tools] || {})[opts[:name].to_s.to_sym]
  return 0 unless t

  Array(t[:prm_window]).length
rescue StandardError
  0
end

.proxy_trustObject

P4 helper — Registry.rank calls this so β·advantage is scaled down when the proxy is untrustworthy.



191
192
193
194
195
196
# File 'lib/pwn/ai/agent/metrics.rb', line 191

public_class_method def self.proxy_trust
  d = defined?(Reward) && Reward.respond_to?(:proxy_distrust) ? Reward.proxy_distrust : 0.0
  (1.0 - d.to_f).clamp(0.0, 1.0)
rescue StandardError
  1.0
end

.record(opts = {}) ⇒ Object

Supported Method Parameters

PWN::AI::Agent::Metrics.record( name: 'required - tool name that was dispatched', success: 'required - Boolean, did the handler complete without error', duration: 'optional - Float seconds the dispatch took', error: 'optional - String error message when success is false', engine: 'optional - Symbol/String AI engine that chose this tool (segments telemetry)' )



83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
# File 'lib/pwn/ai/agent/metrics.rb', line 83

public_class_method def self.record(opts = {})
  name     = opts[:name].to_s
  success  = opts[:success] ? true : false
  duration = opts[:duration].to_f
  error    = opts[:error]
  engine   = opts[:engine].to_s
  return if name.empty?

  metrics = load
  metrics[:tools] ||= {}
  key = name.to_sym
  t = metrics[:tools][key] ||= blank_bucket
  bump(bucket: t, success: success, duration: duration, error: error)
  unless engine.empty?
    t[:engines] ||= {}
    e = t[:engines][engine.to_sym] ||= blank_bucket
    bump(bucket: e, success: success, duration: duration, error: error)
  end
  save(metrics: metrics)
  t
end

.record_calibration(opts = {}) ⇒ Object

Supported Method Parameters

PWN::AI::Agent::Metrics.record_calibration(predicted:, actual:, brier:, engine:)

W3 — plan_first emits p(success); Loop.run calls this with the realised outcome. Tracked per-engine so calibration of the local LoRA vs frontier is comparable.



447
448
449
450
451
452
453
454
455
456
457
458
# File 'lib/pwn/ai/agent/metrics.rb', line 447

public_class_method def self.record_calibration(opts = {})
  m = load
  m[:calibration] ||= {}
  eng = (opts[:engine] || :global).to_s.to_sym
  c = m[:calibration][eng] ||= { n: 0, brier_sum: 0.0, p_sum: 0.0, a_sum: 0.0 }
  c[:n]         += 1
  c[:brier_sum] += opts[:brier].to_f
  c[:p_sum]     += opts[:predicted].to_f
  c[:a_sum]     += opts[:actual].to_f
  save(metrics: m)
  c
end

.record_judge(opts = {}) ⇒ Object

P20 — fold ORM judge (0..1) into per-tool telemetry so UCB / Thompson / advantage can prefer judge-grounded rates over the inflated handler-ok proxy when proxy_distrust is high.



326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
# File 'lib/pwn/ai/agent/metrics.rb', line 326

public_class_method def self.record_judge(opts = {})
  name = opts[:name].to_s
  return if name.empty?

  score = opts[:score].to_f.clamp(0.0, 1.0)
  conf  = opts.key?(:confidence) ? opts[:confidence].to_f.clamp(0.0, 1.0) : 0.7
  src   = opts[:source].to_s
  src   = 'heuristic' if src.empty?
  m = load
  m[:tools] ||= {}
  t = m[:tools][name.to_sym] ||= blank_bucket
  t[:judge_window] = (Array(t[:judge_window]) + [score]).last(40)
  t[:judge_conf_window] = (Array(t[:judge_conf_window]) + [conf]).last(40)
  t[:judge_src_window] = (Array(t[:judge_src_window]) + [src]).last(40)
  t[:judge_sum] = t[:judge_window].sum.to_f
  t[:judge_n] = t[:judge_window].length
  save(metrics: m)
  t
rescue StandardError
  nil
end

.record_step_reward(opts = {}) ⇒ Object

P18 — called by Reward.prm after session annotate so live routing can bias toward tools that recently advanced goals (+1 step_reward).



306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
# File 'lib/pwn/ai/agent/metrics.rb', line 306

public_class_method def self.record_step_reward(opts = {})
  name = opts[:name].to_s
  return if name.empty?

  rew = opts[:reward].to_f.clamp(-1.0, 1.0)
  m = load
  m[:tools] ||= {}
  t = m[:tools][name.to_sym] ||= blank_bucket
  t[:prm_window] = (Array(t[:prm_window]) + [rew]).last(40)
  t[:prm_sum] = t[:prm_window].sum
  t[:prm_n] = t[:prm_window].length
  save(metrics: m)
  t
rescue StandardError
  nil
end

.resetObject



506
507
508
509
# File 'lib/pwn/ai/agent/metrics.rb', line 506

public_class_method def self.reset
  FileUtils.rm_f(METRICS_FILE)
  { tools: {}, updated_at: nil }
end

.save(opts = {}) ⇒ Object

Supported Method Parameters

PWN::AI::Agent::Metrics.save( metrics: 'required - Hash returned by .load / mutated in place' )



54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
# File 'lib/pwn/ai/agent/metrics.rb', line 54

public_class_method def self.save(opts = {})
  metrics = opts[:metrics] ||= { tools: {} }
  metrics[:updated_at] = Time.now.utc.iso8601
  FileUtils.mkdir_p(File.dirname(METRICS_FILE))
  # 4.4 — flock + atomic rename
  path = METRICS_FILE
  tmp  = File.join(File.dirname(path), ".#{File.basename(path)}.#{Process.pid}.tmp")
  body = JSON.pretty_generate(metrics)
  File.open(tmp, File::WRONLY | File::CREAT | File::TRUNC, 0o644) do |f|
    f.flock(File::LOCK_EX)
    f.write(body)
    f.flush
    f.fsync
  end
  File.rename(tmp, path)
  metrics
ensure
  FileUtils.rm_f(tmp) if defined?(tmp) && tmp && File.exist?(tmp)
end

.summary(opts = {}) ⇒ Object

Supported Method Parameters

rows = PWN::AI::Agent::Metrics.summary( limit: 'optional - cap number of tools returned (default 25)', engine: 'optional - only that engine's sub-bucket (falls back to global when absent)' )



111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
# File 'lib/pwn/ai/agent/metrics.rb', line 111

public_class_method def self.summary(opts = {})
  limit  = opts[:limit] || 25
  engine = opts[:engine].to_s
  tools  = load[:tools] || {}
  rows = tools.map do |name, t|
    b = engine.empty? ? t : (t.dig(:engines, engine.to_sym) || t)
    calls = b[:calls].to_i
    ok    = b[:ok].to_i
    rate  = calls.positive? ? (ok.to_f / calls).round(3) : 0.0
    avg   = calls.positive? ? (b[:total_duration].to_f / calls).round(3) : 0.0
    {
      name: name.to_s,
      calls: calls,
      success_rate: rate,
      judge_rate: judge_rate(name: name),
      effective_rate: effective_rate(name: name),
      avg_duration: avg,
      last_error: b[:last_error],
      last_at: b[:last_at]
    }
  end
  rows.reject { |r| r[:calls].zero? }
      .sort_by { |r| [-r[:calls], -r[:success_rate]] }.first(limit)
end

.thompson(opts = {}) ⇒ Object

Supported Method Parameters

p = PWN::AI::Agent::Metrics.thompson(name: 'shell')

C1 — Thompson sample from Beta(ok+1, fail+1). Naturally balances exploit/explore; used by Registry.rank as the tie-breaker.



218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
# File 'lib/pwn/ai/agent/metrics.rb', line 218

public_class_method def self.thompson(opts = {})
  t = (load[:tools] || {})[opts[:name].to_s.to_sym] || blank_bucket
  # P20 — when judge samples exist and distrust > 0, tilt Beta toward
  # judge_rate so Thompson explore/exploit tracks ORM not handler-ok.
  ok = t[:ok].to_f
  fail = t[:fail].to_f
  jn = Array(t[:judge_window]).length
  if jn >= 3 && proxy_trust < 0.95
    jr = judge_rate(name: opts[:name]).to_f
    # pseudo-counts from judge window, mixed by distrust
    d = (1.0 - proxy_trust).clamp(0.0, 1.0)
    jok = jr * jn
    jfail = (1.0 - jr) * jn
    ok = (ok * (1.0 - d)) + (jok * d)
    fail = (fail * (1.0 - d)) + (jfail * d)
  end
  beta_sample(alpha: ok + 1.0, beta: fail + 1.0)
rescue StandardError
  0.5
end

.to_context(opts = {}) ⇒ Object

Supported Method Parameters

ctx = PWN::AI::Agent::Metrics.to_context( limit: 'optional - cap number of tools included (default 8)', engine: 'optional - restrict to one engine's telemetry' )



142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
# File 'lib/pwn/ai/agent/metrics.rb', line 142

public_class_method def self.to_context(opts = {})
  limit  = opts[:limit] || 8
  engine = opts[:engine]
  rows   = summary(limit: limit, engine: engine)
  return '' if rows.empty?

  # P4 — when Reward.sentinel says proxy is hacked, haircut displayed
  # success rates so the model does not trust the lie in the prompt.
  distrust = defined?(Reward) && Reward.respond_to?(:proxy_distrust) ? Reward.proxy_distrust : 0.0
  scope = engine.to_s.empty? ? 'historical' : "engine=#{engine}"
  scope = "#{scope}, proxy_distrust=#{distrust.round(2)}" if distrust.positive?
  lines = rows.map do |r|
    # P20 — display effective_rate (judge-blended) when available;
    # fall back to distrust haircut on raw proxy.
    rate = r[:success_rate].to_f
    eff  = r[:effective_rate]
    adj  = if eff
             eff.to_f
           else
             rate - ((rate - 0.5) * distrust)
           end
    err = r[:last_error] ? " last_err=#{r[:last_error][0, 60]}" : ''
    jtag = r[:judge_rate] ? " judge=#{(r[:judge_rate].to_f * 100).round(1)}%" : ''
    tag = distrust.positive? || r[:judge_rate] ? ' (adj)' : ''
    "  - #{r[:name]}: calls=#{r[:calls]} success=#{(adj * 100).round(1)}%#{tag}#{jtag} avg=#{r[:avg_duration]}s#{err}"
  end
  warn_line = distrust.positive? ? "WARNING: reward proxy diverges from judge — success rates haircut by distrust=#{distrust.round(2)}; prefer judge-scored exemplars over raw rates.\n" : ''
  # P0 ops — surface W1 generator_mix when diet is unhealthy so the
  # online controller (and the model) prefer underfilled sources.
  mix_line = ''
  if defined?(Reward) && Reward.respond_to?(:generator_mix)
    begin
      m = Reward.generator_mix
      unless m[:healthy]
        mix_line = "W1 MIX: n=#{m[:n]} traj=#{m[:trajectory_fraction]} " \
                   "urgent=#{Array(m[:urgent]).join(',')} " \
                   "suppress=#{Array(m[:suppress]).join(',')} " \
                   "rec=#{m[:recommendation]}\n"
      end
    rescue StandardError
      mix_line = ''
    end
  end
  health = health_line
  "#{warn_line}#{mix_line}#{health}TOOL EFFECTIVENESS (#{scope}, adapt tool choice accordingly)\n#{lines.join("\n")}\n\n"
end

.ucb(opts = {}) ⇒ Object



198
199
200
201
202
203
204
205
206
207
208
209
210
# File 'lib/pwn/ai/agent/metrics.rb', line 198

public_class_method def self.ucb(opts = {})
  name = opts[:name].to_s
  c    = (opts[:c] || 1.4).to_f
  data = load[:tools] || {}
  t    = data[name.to_sym] || blank_bucket
  n    = [t[:calls].to_f, 1.0].max
  total = [data.values.sum { |v| v[:calls].to_f }, 1.0].max
  # P20 — mean from effective_rate (judge-blended when distrust high)
  mean  = effective_rate(name: name)
  mean + (c * Math.sqrt(Math.log(total) / n))
rescue StandardError
  1.0
end