Module: Vangrail::Providers::Llmlite
- Defined in:
- lib/vangrail/providers/llmlite.rb
Overview
llmlite: a local OpenAI-compatible proxy, and the default this gem builds around.
It is the right default for guardrails specifically. Rails run on every turn, so their latency and their failure modes are the application's; a local endpoint keeps both on this machine, needs no shared credential, and cannot bill anyone. It also means a laptop with the proxy running, and a model named, has working rails with no shared credential to configure.
The proxy serves an instruct model, not a safety classifier, so
model(:guard) is nil and the builder puts a policy rail on the input
side. That is a real difference between endpoints, and it belongs here
rather than in a rail deciding what it is talking to.
Constant Summary collapse
- HOST =
'127.0.0.1'- DEFAULT_PORT =
8760- LISTEN_ERRORS =
[ Errno::ECONNREFUSED, Errno::EHOSTUNREACH, Errno::ENETUNREACH, Errno::ECONNRESET, Errno::ETIMEDOUT, Errno::EADDRNOTAVAIL, SocketError ].freeze
Class Method Summary collapse
- .base_url(env = ENV) ⇒ Object
-
.embed_model(env = ENV) ⇒ Object
No default.
- .host(env = ENV) ⇒ Object
- .key(env = ENV) ⇒ Object
-
.listening?(env = ENV) ⇒ Boolean
A TCP connect, not a request: a proxy that is not running is the common case, and finding that out must cost microseconds rather than a timeout.
- .model(env = ENV) ⇒ Object
- .port(env = ENV) ⇒ Object
- .present(value) ⇒ Object
- .provider(env = ENV) ⇒ Object
Class Method Details
.base_url(env = ENV) ⇒ Object
39 40 41 |
# File 'lib/vangrail/providers/llmlite.rb', line 39 def base_url(env = ENV) "http://#{host(env)}:#{port(env)}/v1" end |
.embed_model(env = ENV) ⇒ Object
No default. Which embedding model a proxy serves, if any, is deployment knowledge, and a guessed name costs a 404 on every check while looking like a rail that ran.
60 61 62 |
# File 'lib/vangrail/providers/llmlite.rb', line 60 def (env = ENV) present(env['LLMLITE_EMBED_MODEL'] || env['GUARDRAILS_EMBED_MODEL']) end |
.host(env = ENV) ⇒ Object
35 36 37 |
# File 'lib/vangrail/providers/llmlite.rb', line 35 def host(env = ENV) env['LLMLITE_HOST'].to_s.strip.empty? ? HOST : env['LLMLITE_HOST'].strip end |
.key(env = ENV) ⇒ Object
64 65 66 |
# File 'lib/vangrail/providers/llmlite.rb', line 64 def key(env = ENV) present(env['LLMLITE_API_KEY']) end |
.listening?(env = ENV) ⇒ Boolean
A TCP connect, not a request: a proxy that is not running is the common case, and finding that out must cost microseconds rather than a timeout.
45 46 47 48 49 50 51 |
# File 'lib/vangrail/providers/llmlite.rb', line 45 def listening?(env = ENV) socket = TCPSocket.new(host(env), port(env)) socket.close true rescue *LISTEN_ERRORS false end |
.model(env = ENV) ⇒ Object
53 54 55 |
# File 'lib/vangrail/providers/llmlite.rb', line 53 def model(env = ENV) present(env['LLMLITE_MODEL'] || env['GROK_LLMLITE_MODEL']) end |
.port(env = ENV) ⇒ Object
31 32 33 |
# File 'lib/vangrail/providers/llmlite.rb', line 31 def port(env = ENV) (env['LLMLITE_PORT'] || env['GROK_SHIM_PORT'] || DEFAULT_PORT).to_i end |
.present(value) ⇒ Object
80 81 82 83 |
# File 'lib/vangrail/providers/llmlite.rb', line 80 def present(value) s = value.to_s.strip s.empty? ? nil : s end |
.provider(env = ENV) ⇒ Object
68 69 70 71 72 73 74 75 76 77 78 |
# File 'lib/vangrail/providers/llmlite.rb', line 68 def provider(env = ENV) resolved = key(env) Provider.new( name: 'llmlite', base_url: base_url(env), models: { judge: model(env), guard: nil, embed: (env) }, key_resolver: resolved && -> { resolved }, local: true, probe: -> { listening?(env) }, ) end |