Module: PWN::AI::Ollama
- Defined in:
- lib/pwn/ai/ollama.rb
Overview
This plugin is used for interacting w/ Ollama's REST API using the 'rest' browser type of PWN::Plugins::TransparentBrowser. This is based on the following Ollama API Specification: https://api.openai.com/v1
Class Method Summary collapse
-
.authors ⇒ Object
- Author(s)
0day Inc.
-
.chat(opts = {}) ⇒ Object
- Supported Method Parameters
response = PWN::AI::Ollama.chat( request: 'required - message to Ollama' model: 'optional - model to use for text generation (defaults to PWN::Env[:ollama][:model])', temp: 'optional - creative response float (deafults to PWN::Env[:ollama][:temp])', system_role_content: 'optional - context to set up the model behavior for conversation (Default: PWN::Env[:ollama][:system_role_content])', response_history: 'optional - pass response back in to have a conversation', speak_answer: 'optional speak answer using PWN::Plugins::Voice.text_to_speech (Default: nil)', timeout: 'optional timeout in seconds (defaults to 900)', spinner: 'optional - display spinner (defaults to false)' ).
-
.chat_with_tools(opts = {}) ⇒ Object
- Supported Method Parameters
response = PWN::AI::Ollama.chat_with_tools( messages: 'required - full OpenAI-format messages array (system/user/assistant/tool)', tools: 'optional - OpenAI tools array [function:{...}]', tool_choice: 'optional - "auto" | "none" | function:{name:..}', model: 'optional - overrides PWN::Env[:ollama][:model]', temp: 'optional - temperature (defaults to PWN::Env[:ollama][:temp] || 1)', timeout: 'optional - seconds (default 900)', spinner: 'optional - display spinner (default false)' ).
-
.get_models ⇒ Object
- Supported Method Parameters
response = PWN::AI::Ollama.get_models.
-
.help ⇒ Object
Display Usage for this Module.
Class Method Details
.authors ⇒ Object
- Author(s)
0day Inc. support@0dayinc.com
577 578 579 580 581 |
# File 'lib/pwn/ai/ollama.rb', line 577 public_class_method def self. "AUTHOR(S): 0day Inc. <support@0dayinc.com> " end |
.chat(opts = {}) ⇒ Object
- Supported Method Parameters
response = PWN::AI::Ollama.chat( request: 'required - message to Ollama' model: 'optional - model to use for text generation (defaults to PWN::Env[:ollama][:model])', temp: 'optional - creative response float (deafults to PWN::Env[:ollama][:temp])', system_role_content: 'optional - context to set up the model behavior for conversation (Default: PWN::Env[:ollama][:system_role_content])', response_history: 'optional - pass response back in to have a conversation', speak_answer: 'optional speak answer using PWN::Plugins::Voice.text_to_speech (Default: nil)', timeout: 'optional timeout in seconds (defaults to 900)', spinner: 'optional - display spinner (defaults to false)' )
493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 |
# File 'lib/pwn/ai/ollama.rb', line 493 public_class_method def self.chat(opts = {}) engine = PWN::Env[:ai][:ollama] request = opts[:request] max_prompt_length = engine[:max_prompt_length] ||= 1_000_000 request_trunc_idx = ((max_prompt_length - 1) / 3.36).floor request = request[0..request_trunc_idx] model = opts[:model] ||= engine[:model] raise 'ERROR: Model is required. Call #get_models method for details' if model.nil? temp = opts[:temp].to_f ||= engine[:temp].to_f temp = 1 if temp.zero? rest_call = 'ollama/v1/chat/completions' response_history = opts[:response_history] max_tokens = response_history[:usage][:total_tokens] unless response_history.nil? system_role_content = opts[:system_role_content] ||= engine[:system_role_content] system_role = { role: 'system', content: system_role_content } user_role = { role: 'user', content: request } response_history ||= { choices: [system_role] } choices_len = response_history[:choices].length http_body = { model: model, messages: [system_role], temperature: temp, stream: true } if response_history[:choices].length > 1 response_history[:choices][1..-1].each do || http_body[:messages].push() end end http_body[:messages].push(user_role) timeout = opts[:timeout] spinner = opts[:spinner] response = ollama_rest_call( http_method: :post, rest_call: rest_call, http_body: http_body, timeout: timeout, spinner: spinner ) json_resp = JSON.parse(response, symbolize_names: true) assistant_resp = json_resp[:choices].first[:message] json_resp[:choices] = http_body[:messages] json_resp[:choices].push(assistant_resp) speak_answer = true if opts[:speak_answer] if speak_answer answer = assistant_resp[:content] text_path = "/tmp/#{SecureRandom.hex}.pwn_voice" # answer = json_resp[:choices].last[:text] # answer = json_resp[:choices].last[:content] if gpt File.write(text_path, answer) PWN::Plugins::Voice.text_to_speech(text_path: text_path) File.unlink(text_path) end json_resp rescue StandardError => e raise e end |
.chat_with_tools(opts = {}) ⇒ Object
- Supported Method Parameters
response = PWN::AI::Ollama.chat_with_tools( messages: 'required - full OpenAI-format messages array (system/user/assistant/tool)', tools: 'optional - OpenAI tools array [function:{...}]', tool_choice: 'optional - "auto" | "none" | function:{name:..}', model: 'optional - overrides PWN::Env[:ollama][:model]', temp: 'optional - temperature (defaults to PWN::Env[:ollama][:temp] || 1)', timeout: 'optional - seconds (default 900)', spinner: 'optional - display spinner (default false)' )
Returns a Hash with :choices / :assistant_message intact (including :message) — used by PWN::AI::Agent::Loop.
LOCAL-MODEL SCAFFOLDING
This hits Ollama's NATIVE /api/chat (not the OpenAI-compat shim) so the following actually take effect:
options.num_ctx - Ollama defaults to 2048; the pwn-ai system
prompt alone blows that. Defaults here to
PWN::Env[:ai][:ollama][:num_ctx] || 32768.
options.temperature - forced to 0.1 on tool-bearing turns for
deterministic tool selection; engine[:temp]
(creative) on the final text-only turn.
format - ONLY set when engine[:format] is explicitly
configured. Never default to 'json' when
tools: are present — that fights native
tool_calls and kills mid-loop tool use.
keep_alive: '30m' - avoids reload latency between iterations.
tool_calls come back with function.arguments as a Hash (not a JSON string), which PWN::AI::Agent::Dispatch.parse_args handles. Streaming is ON (stream: true): ollama_rest_call assembles NDJSON chunks back into a single response so the return shape is unchanged.
411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 |
# File 'lib/pwn/ai/ollama.rb', line 411 public_class_method def self.chat_with_tools(opts = {}) engine = PWN::Env[:ai][:ollama] = opts[:messages] raise 'ERROR: messages array is required' if .nil? || .empty? model = opts[:model] ||= engine[:model] raise 'ERROR: Model is required. Call #get_models method for details' if model.nil? temp = opts[:temp].to_f temp = engine[:temp].to_f.nonzero? || 1 if temp.zero? tools_present = opts[:tools] && !opts[:tools].empty? tool_temp = (engine[:tool_temp] || 0.1).to_f num_ctx = (engine[:num_ctx] || 32_768).to_i # Hard cap generation length. Thinking models (Qwen3 / R1-style) # otherwise stream unbounded message.thinking tokens until the # idle read_timeout (default 900s) — which looks like a 15–25 min # "stuck spinner" while bytes keep arriving on the socket. num_predict = (engine[:num_predict] || 4_096).to_i keep_alive = engine[:keep_alive] || '30m' http_body = { model: model, messages: , stream: true, keep_alive: keep_alive, options: { num_ctx: num_ctx, num_predict: num_predict, temperature: tools_present ? tool_temp : temp } } if tools_present http_body[:tools] = opts[:tools] # 0.2 — omit format when tools are present unless the operator # explicitly set PWN::Env[:ai][:ollama][:format]. Forcing 'json' # races the native tool_calls sampler and produces "can't call tools" # mid-loop on many local models. fmt = engine[:format] http_body[:format] = fmt unless fmt.nil? || fmt.to_s.empty? end http_body[:tool_choice] = opts[:tool_choice] if opts[:tool_choice] response = ollama_rest_call( http_method: :post, rest_call: 'ollama/api/chat', http_body: http_body, timeout: opts[:timeout], spinner: opts[:spinner] ) raise 'ERROR: Ollama chat_with_tools received empty response from ollama_rest_call' if response.nil? || (response.respond_to?(:empty?) && response.empty?) json_resp = JSON.parse(response, symbolize_names: true) # Normalise native /api/chat shape to what Loop.normalize_llm expects. msg = json_resp[:message] || json_resp.dig(:choices, 0, :message) if msg.is_a?(Hash) content = msg[:content].to_s thinking = msg[:thinking].to_s tcalls = Array(msg[:tool_calls]) msg = msg.merge(content: visible_from_thinking(thinking: thinking)) if content.strip.empty? && !thinking.strip.empty? && tcalls.empty? end json_resp[:choices] = [{ message: msg }] if msg && !json_resp.key?(:choices) json_resp[:assistant_message] = msg raise "ERROR: Ollama response missing message/choices: #{json_resp.inspect[0, 400]}" if msg.nil? json_resp rescue StandardError => e raise e end |
.get_models ⇒ Object
- Supported Method Parameters
response = PWN::AI::Ollama.get_models
369 370 371 372 373 374 375 |
# File 'lib/pwn/ai/ollama.rb', line 369 public_class_method def self.get_models models = ollama_rest_call(rest_call: 'ollama/api/tags') JSON.parse(models, symbolize_names: true)[:models] rescue StandardError => e raise e end |
.help ⇒ Object
Display Usage for this Module
585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 |
# File 'lib/pwn/ai/ollama.rb', line 585 public_class_method def self.help puts "USAGE: models = #{self}.get_models response = #{self}.chat( request: 'required - message to Ollama', model: 'optional - model to use for text generation (defaults to PWN::Env[:ai][:ollama][:model])', temp: 'optional - creative response float (defaults to PWN::Env[:ai][:ollama][:temp])', system_role_content: 'optional - context to set up the model behavior for conversation (Default: PWN::Env[:ai][:ollama][:system_role_content])', response_history: 'optional - pass response back in to have a conversation', speak_answer: 'optional speak answer using PWN::Plugins::Voice.text_to_speech (Default: nil)', timeout: 'optional - timeout in seconds (defaults to 900)', spinner: 'optional - display spinner (defaults to false)' ) #{self}.authors " end |