Traject::SolrPool
Traject::SolrPool::SolrJsonWriter is a drop-in pooled replacement for the
stock Traject::SolrJsonWriter. It routes every Solr update through a
persistent, credential-isolated, thread-safe HTTP connection pool provided by
http_connection_pool instead of opening a fresh HTTPClient
connection per request.
It reuses the stock writer's solr.* / solr_writer.* settings vocabulary, so
it is a drop-in wherever those settings are practical. The writer is:
- thread-safe (shared state uses
concurrent-rubyprimitives; a borrowed connection is never held across threads), - usable from Sidekiq >= 8 workers (fork-safe via
connection_pool >= 2.5), - Zeitwerk-conformant, and
- compatible with Rails (no Rails runtime dependency).
All four properties are verified by the test suite (unit, concurrency, Sidekiq, Zeitwerk, and Rails-coexistence specs).
Installation
Add the gem to your application's Gemfile:
gem 'traject-solr_pool'
Then install:
bundle install
Requirements
- Ruby >= 3.3
traject>= 3.9 (3.9.0 relaxed its HTTP cap tohttp >= 3.0, < 7, so it co-installs withhttp_connection_pool'shttp ~> 6.0directly from RubyGems — no edge checkout needed)http_connection_pool>= 0.1.1
Earlier releases of traject capped their HTTP dependency at http < 6, which
collided with http_connection_pool's http ~> 6.0; the Gemfile once carried a
path: override to a local edge checkout to work around that. It is no longer
required and the override is commented out.
Usage
Register the writer with traject's writer_class_name setting:
require 'traject/solr_pool'
settings do
provide 'writer_class_name', 'Traject::SolrPool::SolrJsonWriter'
provide 'solr.url', 'http://localhost:8983/solr/my_core'
provide 'solr_writer.thread_pool', 4
provide 'solr_pool.pool_size', 5
end
The origin (scheme://host:port) derived from solr.url becomes the pool's
base_url; the remainder of the URL becomes the relative request path. Basic
auth credentials are sent as an Authorization header, never baked into the
origin, so distinct credentials resolve to distinct pools and never share
connections.
Settings
The writer honours the stock solr.* / solr_writer.* vocabulary, plus one
new namespaced setting.
| Setting | Default | Role |
|---|---|---|
solr.url |
(required unless solr.update_url set) |
Base Solr URL; /update/json is derived from it |
solr.update_url |
derived from solr.url |
Full update handler URL; used verbatim when provided |
solr_writer.batch_size |
100 |
Documents per batched update |
solr_writer.thread_pool |
1 |
Number of background writer threads |
solr_writer.max_skipped |
0 |
Skip tolerance before MaxSkippedRecordsExceeded (negative disables the cap) |
solr_writer.skippable_exceptions |
http.rb-native list (see below) | Exceptions treated as skippable during per-record retry |
solr_writer.commit_on_close |
false |
Send a commit when the writer closes (legacy solrj_writer.commit_on_close also honoured) |
solr_writer.solr_update_args |
none | Query params (e.g. { 'commitWithin' => 1000 }) applied to every update and delete request |
solr_writer.commit_solr_update_args |
{ 'commit' => 'true' } |
Query params for the commit request |
solr_writer.commit_timeout |
600 (10 min) |
Read timeout applied to the commit request only, so a slow commit is not cut off by a short http_timeout |
solr_writer.basic_auth_user |
embedded URI user | Basic-auth user (overrides credentials embedded in the URL) |
solr_writer.basic_auth_password |
embedded URI password | Basic-auth password (overrides credentials embedded in the URL) |
solr_writer.http_timeout |
none | Per-request HTTP timeout, passed to the pool connection |
solr_writer.pool_timeout |
pool default | Max time to wait for a connection checkout from the pool |
solr_pool.pool_size |
solr_writer.thread_pool + 1 |
Pool capacity, sized so no writer thread starves on checkout while the caller thread can still flush/commit |
Dropped settings
Two solr_json_writer.* settings from the stock writer are HTTPClient-specific
and are not supported:
solr_json_writer.http_client— the pool owns connection construction, so there is noHTTPClientinstance to inject.solr_json_writer.use_packaged_certs— anHTTPClientssl_config quirk; http.rb uses the operating system's certificate store instead.
Error handling
-
Traject::SolrPool::SolrJsonWriter::BadHttpResponse(aRuntimeError) is raised on any non-200 response. Its#responseaccessor returns the raw http.rb response (use#code/#to_s), and it extracts Solr's JSONerror.msginto the message when present. -
The default
solr_writer.skippable_exceptionslist is http.rb-native:[ HTTP::TimeoutError, HttpConnectionPool::TimeoutError, HTTP::ConnectionError, SocketError, Errno::ECONNREFUSED, Traject::SolrPool::SolrJsonWriter::BadHttpResponse ]A batch that fails is retried record by record; a record that raises a skippable exception is logged and counted, aborting the run only once
solr_writer.max_skippedis exceeded. -
Setting
solr_writer.skippable_exceptionsexplicitly replaces the default list entirely with your own. -
commit,delete, anddelete_all!raise raw on failure; they are not part of the skip path. -
Credentials never appear in logs or error messages: auth lives in a header and logs show the origin only.
Pool lifecycle
close flushes any queued records, waits for the background threads, and
commits when solr_writer.commit_on_close is set. It then leaves the pool
warm in the http_connection_pool registry so later writers (or a future
reader) on the same origin reuse it.
Teardown is the host application's concern. Close the pools explicitly on process shutdown or after a fork, for example:
HttpConnectionPool::Registry.instance.close_all
Development
Run the full CI pipeline (offline bundler-audit, RuboCop, RSpec):
bundle exec rake ci
Individual tasks (rake spec, rake rubocop, rake bundle:audit:check) are
also available.
Contributing
Bug reports and pull requests are welcome. Releasing and pushing are maintainer responsibilities — contributors should not push tags or publish gems.
License
The gem is available as open source under the terms of the MIT License.