Module: PWN::AI::Agent::Loop
- Defined in:
- lib/pwn/ai/agent/loop.rb
Overview
The agent conversation loop:
build system prompt → call LLM with tools → if tool_calls: dispatch,
append role:'tool' results, loop → else: return text.
This replaces the regex-ReAct in PWN::Plugins::REPL :pwn_ai_hook with native function-calling. State (memory, skills, sessions) is all externalised — Loop.run is stateless aside from the messages array it builds.
NEGATIVE-FEEDBACK CLOSURE
Loop.run is where "learn from mistakes, don't repeat them" is actually enforced. On EVERY failed dispatch it:
1. Records the (tool, normalised_error) fingerprint into
PWN::AI::Agent::Mistakes with a PERSISTENT cross-session count.
2. Reads that count back and, if it OR the in-turn count reaches
REPEAT_THRESHOLD, prepends a hard "REPEATED FAILURE — change
approach" guard to the tool result the model sees next.
3. Appends Mistakes.correction_hint (seen N×, sig, KNOWN FIX: …)
so a previously-discovered fix is handed straight back to the
model on the FIRST recurrence in a new session — it does not
have to fail 3× again to re-learn what it already knew.
PromptBuilder.mistakes_block re-injects the top open mistakes and top known fixes into the system prompt of every future turn.
LOCAL-MODEL SCAFFOLDING
When the active engine is :ollama (or the corresponding :agent flags are set) Loop.run additionally:
* threads request → PromptBuilder for relevance-ranked MEMORY,
* threads request → Registry.definitions(relevance:) for a slimmed
tool set (:tool_router),
* splices Learning.exemplars_for(request:) between system and user
as few-shot behaviour retrieval,
* runs a plan-then-act pre-pass (:plan_first) so the model
externalises a tool plan before its first dispatch,
* escalates to a frontier persona for a 3-line corrective hint
once ≥ ESCALATE_AFTER_FAILS in-turn failures accumulate
(:escalation_persona) — the local model still produces the final
answer so Learning/Metrics stay attributed to :ollama.
Constant Summary collapse
- DEFAULT_MAX_ITERS =
777- ESCALATE_AFTER_FAILS =
4- BUDGET_HARD_STOP_FAILS =
P17 — when empty_final / known thrash shapes dominate, stop before burning the full ollama cap so the corpus is not pure terminal failure.
8- BUDGET_EMPTY_FINAL_STOP =
3- ENGINE_MODS =
{ openai: 'PWN::AI::OpenAI', grok: 'PWN::AI::Grok', ollama: 'PWN::AI::Ollama', anthropic: 'PWN::AI::Anthropic', gemini: 'PWN::AI::Gemini' }.freeze
- INCOMPLETE_FINAL_RX =
P28 — incomplete / handoff finals: model emitted text-only before the goal was done ("shall I proceed?", "next step:", "want me to…"). Loop.run treats no-tool_calls as FINAL; this detector lets us refuse that handoff and keep the tool loop alive for multi-step autonomy.
/ \b(shall\s+i|should\s+i|may\s+i|can\s+i|want\s+me\s+to|do\s+you\s+want\s+me| next\s+single\s+step|next\s+step\s*:|awaiting\s+your\s+(ok|approval|go-ahead|confirmation)| if\s+you(?:'d|\s+would)\s+like\s+me\s+to|say\s+the\s+word|confirm\s+(before|and\s+i)| ready\s+to\s+proceed|ok\s+to\s+(proceed|continue|apply)|proceed\?| continue\?|before\s+i\s+(apply|change|run|continue|proceed)| once\s+you\s+(confirm|approve)|let\s+me\s+know\s+if| i(?:'ll|\s+will)\s+wait\b|waiting\s+for\s+(your\s+)?(go|ok|approval|confirmation) )\b /ix
Class Method Summary collapse
-
.authors ⇒ Object
- Author(s)
0day Inc.
-
.help ⇒ Object
Display Usage for this Module.
-
.run(opts = {}) ⇒ Object
- Supported Method Parameters
final = PWN::AI::Agent::Loop.run( request: 'required - what the human typed', session_id: 'optional - PWN::Sessions id (transcript is appended to it)', enabled_toolsets: 'optional - subset of Registry.toolsets, or nil for all', on_tool: 'optional - ->(name, args, result) callback for live UI', system_role_content: 'optional - override default system prompt (built from session_id if not provided)' ).
Class Method Details
.authors ⇒ Object
- Author(s)
0day Inc. support@0dayinc.com
804 805 806 |
# File 'lib/pwn/ai/agent/loop.rb', line 804 public_class_method def self. "AUTHOR(S):\n 0day Inc. <support@0dayinc.com>\n" end |
.help ⇒ Object
Display Usage for this Module
810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 |
# File 'lib/pwn/ai/agent/loop.rb', line 810 public_class_method def self.help puts <<~USAGE USAGE: final = PWN::AI::Agent::Loop.run( request: 'what does `id` return on this host?', session_id: PWN::Sessions.create[:id], enabled_toolsets: %w[terminal pwn memory skills], on_tool: ->(name, args, result) { puts "→ \#{name}: \#{result[0,1_024]}" }, system_role_content: 'You are a helpful assistant that can call tools to answer questions.' ) Supported engines: #{ENGINE_MODS.keys.join(', ')} Set PWN::Env[:ai][:active] to choose; PWN::Env[:ai][:agent][:max_iters] to bound. Local-model scaffolding (PWN::Env[:ai][:agent]): :plan_first - Boolean, plan-then-act pre-pass (default: engine == :ollama) :tool_router - Boolean/nil, slim Registry.definitions (nil=auto on for ollama) :escalation_persona - Swarm persona name for frontier corrective hints when stuck :critic - S3 constitutional critic before every final (Boolean) :red_team_plan - S4 adversarial plan review after plan_first (Boolean) :counterfactual - S2 A/B branch on REPEAT_THRESHOLD → DPO pair (Boolean) :hindsight - C3 HER-relabel failures (Boolean, default true) :verify_as_reward - E3 ground every final via extro_verify (Boolean) P28 autonomy: incomplete-final detector refuses mid-goal handoffs; W3 overconf max_iters_cap is 40 on remote engines (8 on ollama). P17 budget-hot always caps max_iters to 8 for ALL engines (not only ollama). #{self}.authors USAGE end |
.run(opts = {}) ⇒ Object
- Supported Method Parameters
final = PWN::AI::Agent::Loop.run( request: 'required - what the human typed', session_id: 'optional - PWN::Sessions id (transcript is appended to it)', enabled_toolsets: 'optional - subset of Registry.toolsets, or nil for all', on_tool: 'optional - ->(name, args, result) callback for live UI', system_role_content: 'optional - override default system prompt (built from session_id if not provided)' )
583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 |
# File 'lib/pwn/ai/agent/loop.rb', line 583 public_class_method def self.run(opts = {}) request = opts[:request].to_s session_id = opts[:session_id] on_tool = opts[:on_tool] engine = active_engine local = engine == :ollama system_role_content = opts[:system_role_content] ||= PWN::AI::Agent::PromptBuilder.build(session_id: session_id, request: request) Registry.discover expose_current_session(session_id: session_id) Mistakes.check_user_correction(request: request, session_id: session_id) if defined?(Mistakes) tools = Registry.definitions(enabled: opts[:enabled_toolsets], relevance: request) = [{ role: 'system', content: system_role_content }] .concat(Learning.exemplars_for(request: request)) if local && defined?(Learning) && Learning.respond_to?(:exemplars_for) << { role: 'user', content: request } append_session(session_id: session_id, role: 'user', content: request) predicted = nil Thread.current[:pwn_plan_predicted] = nil cal_state = calibration_state force_plan = cal_state[:force_plan] if (force_plan || agent_flag(key: :plan_first, default: local) || budget_exhaustion_hot?) && !Array(tools).empty? predicted = plan_first(messages: , request: request) # P22 — prefer explicit return; fall back to thread stash predicted = Thread.current[:pwn_plan_predicted] if predicted.nil? end if budget_exhaustion_hot? << { role: 'user', content: '[pwn-ai/p17] Budget-exhaustion is the top open failure on this host. Prefer the SHORTEST plan that finishes the ask (≤3 tool calls). Emit a final answer as soon as you have evidence — do not explore.' } end if force_plan && cal_state[:cal] << { role: 'user', content: "[pwn-ai/w3] engine=#{active_engine} is overconfident " \ "(brier=#{cal_state[:cal][:brier]}, overconf=#{cal_state[:cal][:overconfidence]}). " \ 'Prefer high-judge exemplars, verify claims, and avoid speculative tool calls.' } end turn_fails = Hash.new(0) escalated = false max_iters.times do |i| # 3.1 — compact history on local so tool dumps don't fill num_ctx compact_history!(messages: ) if local # P17 — on the final iteration, strip tools and demand a plain-text # answer. Without this the model happily emits one more tool_calls # batch, burns the last slot, and lands on budget_exhausted with # nothing the user (or ORM) can use. # P17 deepen — when budget_hot, force text-only on the LAST TWO # iters so a final tool_calls batch cannot burn the terminal slot. # P17 deepen³ — under hot, force text-only on last THREE of the # 8-iter cap so a late tool binge cannot burn every salvage slot. text_only_iters = budget_exhaustion_hot? ? 3 : 1 last_iter = (i >= max_iters - text_only_iters) if last_iter tag = i >= max_iters - 1 ? 'FINAL ITERATION' : 'PENULTIMATE — wrap up' << { role: 'user', content: "[pwn-ai/p17] #{tag} — do NOT call any more tools. " \ 'Write the best complete answer you can from evidence already in this ' \ 'transcript. If the goal is unfinished, report exactly what is done, ' \ 'what is blocked, and the concrete remaining work — do NOT ask the ' \ 'user to confirm the next step.' } end msg = call_engine(messages: , tools: last_iter ? nil : tools) return '[pwn-ai] engine returned no message' if msg.nil? calls = Array(msg[:tool_calls]) text = msg[:content].to_s # Empty-final guard (local/thinking models): Ollama sometimes # returns done_reason=stop with eval_count<=1, empty content, no # tool_calls — historically surface as a blank TUI reply. Do NOT # commit that as the answer; drop the empty assistant turn, # inject a one-shot nudge, and keep iterating. if calls.empty? && text.strip.empty? warn "[pwn-ai/loop] empty final from #{engine} on iter=#{i}; nudging" if local << { role: 'user', content: 'Your previous reply was empty (no tool_calls and no content). ' \ 'Either call a tool now, or write the final answer for the user as plain text. ' \ 'Do not reply with an empty message.' } turn_fails['empty_final'] += 1 next end << msg if calls.empty? # P28 — refuse polite mid-goal handoffs so multi-step tasks stay autonomous. if incomplete_final?(text: text, last_iter: last_iter) && turn_fails['incomplete_final'].to_i < 2 turn_fails['incomplete_final'] += 1 warn "[pwn-ai/loop] incomplete final on iter=#{i}; continuing autonomously" << { role: 'user', content: '[pwn-ai/p28] That reply handed control back before the goal was done. ' \ 'Do NOT ask the user to confirm the next step. Continue with the ' \ 'necessary tool calls now and finish the goal autonomously. Only ' \ 'emit a final answer when the request is complete or truly blocked.' } next end append_session(session_id: session_id, role: 'assistant', content: text) Learning.auto_introspect(session_id: session_id, request: request, final: text, predicted: predicted) if defined?(Learning) && should_auto_introspect?(local: local, turn_fails: turn_fails, iter: i) return text end calls.each do |tc| name = tc.dig(:function, :name).to_s args = tc.dig(:function, :arguments) entry = Registry.lookup(name: name) started = Time.now raw = Dispatch.call(tool_call: tc) tele = record_metrics(name: name, started: started, raw: raw, args: args, session_id: session_id, engine: engine) result = Result.condition(content: raw, entry: entry) unless tele[:ok] fkey = Digest::SHA256.hexdigest("#{name}|#{args}")[0, 16] turn_fails[fkey] += 1 persist = tele.dig(:mistake, :count).to_i count = [turn_fails[fkey], persist].max hint = defined?(Mistakes) ? Mistakes.correction_hint(tool: name, error: tele[:err] || raw[0, 300]) : '' # S2 — counterfactual A/B: at the repeat threshold, fork an # alt-persona branch, judge both, inject the winner. Real # advantage estimation; (loser, winner) → DPO preference. thresh = defined?(Mistakes) ? Mistakes::REPEAT_THRESHOLD : 3 # P17 — never fork counterfactual when budget fingerprints dominate: # CF is another mini agent loop and is the #1 amplifier of # iteration-budget exhaustion on this host. if count >= thresh && !escalated && defined?(Curriculum) && !budget_exhaustion_hot? cf = (turn_fails["cf:#{fkey}"] += 1) == 1 ? Curriculum.counterfactual(request: request, name: name, args: args, error: tele[:err] || raw[0, 200], hint: hint) : nil hint = "#{hint}\n[pwn-ai/counterfactual] branch #{cf[:branch]} (score=#{cf[:score].round(2)}): #{cf[:content]}" if cf end result = guard_repeated_failure(name: name, count: count, hint: hint, result: result) end on_tool&.call(name, args, result) << { role: 'tool', tool_call_id: tc[:id] || tc['id'] || "call_#{i}", name: name, content: result } append_session( session_id: session_id, role: 'tool', content: "#{name} → #{result[0, 1_024]}" ) end # P17 — hard stop: empty-final thrash or cumulative fails past cap. # Prefer a short apologetic final over another 10 useless tool dumps # that poison ORM/PRM/DPO with terminal failures. empty_n = turn_fails['empty_final'].to_i fail_n = turn_fails.values.sum if empty_n >= BUDGET_EMPTY_FINAL_STOP || fail_n >= BUDGET_HARD_STOP_FAILS msg = if empty_n >= BUDGET_EMPTY_FINAL_STOP '[pwn-ai] stopped: repeated empty finals (budget thrash guard)' else '[pwn-ai] stopped: too many in-turn failures (budget thrash guard)' end if defined?(Mistakes) Mistakes.record( tool: 'agent_loop', error: "budget thrash guard fired empty=#{empty_n} fails=#{fail_n} iter=#{i}", session_id: session_id, source: :loop, shape: :budget_thrash ) end append_session(session_id: session_id, role: 'assistant', content: msg) Learning.auto_introspect(session_id: session_id, request: request, final: msg, predicted: predicted) if defined?(Learning) && should_auto_introspect?(local: local, turn_fails: turn_fails, iter: i) return msg end next unless local && !escalated && turn_fails.values.sum >= ESCALATE_AFTER_FAILS hint = escalate(request: request, turn_fails: turn_fails, session_id: session_id) if hint << { role: 'tool', tool_call_id: "escalation_#{i}", name: 'frontier_hint', content: hint } append_session(session_id: session_id, role: 'tool', content: "frontier_hint → #{hint[0, 1_024]}") end escalated = true end # P17 — exhaust path must still feed Learning so ORM/PRM/HER see the # failure (previously we only Mistakes.record'd and returned a bare # string — no session row, no judge, no hindsight). final_msg = '[pwn-ai] iteration budget exhausted' if defined?(Mistakes) Mistakes.record( tool: 'agent_loop', error: 'iteration budget exhausted without a final answer', session_id: session_id, source: :loop, shape: :budget_exhausted ) end append_session(session_id: session_id, role: 'assistant', content: final_msg) if defined?(Learning) && should_auto_introspect?(local: local, turn_fails: turn_fails, iter: max_iters) Learning.auto_introspect( session_id: session_id, request: request, final: final_msg, predicted: predicted ) end final_msg end |