Traject::SolrPool
Traject::SolrPool::SolrJsonWriter is a drop-in pooled replacement for the
stock Traject::SolrJsonWriter. It routes every Solr update through a
persistent, credential-isolated, thread-safe HTTP connection pool provided by
http_connection_pool instead of opening a fresh HTTPClient
connection per request.
It reuses the stock writer's solr.* / solr_writer.* settings vocabulary, so
it is a drop-in wherever those settings are practical. The writer is:
- thread-safe (shared state uses
concurrent-rubyprimitives; a borrowed connection is never held across threads), - usable from Sidekiq >= 8 workers (fork-safe via
connection_pool >= 2.5), - Zeitwerk-conformant, and
- compatible with Rails (no Rails runtime dependency).
All four properties are verified by the test suite (unit, concurrency, Sidekiq, Zeitwerk, and Rails-coexistence specs).
Installation
Add the gem to your application's Gemfile:
gem 'traject-solr_pool'
Then install:
bundle install
Requirements
- Ruby >= 3.3
traject>= 3.9 (3.9.0 relaxed its HTTP cap tohttp >= 3.0, < 7, so it co-installs withhttp_connection_pool'shttp ~> 6.0directly from RubyGems — no edge checkout needed)http_connection_pool>= 0.1.1
Earlier releases of traject capped their HTTP dependency at http < 6, which
collided with http_connection_pool's http ~> 6.0; the Gemfile once carried a
path: override to a local edge checkout to work around that. It is no longer
required and the override is commented out.
Usage
Register the writer with traject's writer_class_name setting:
require 'traject/solr_pool'
settings do
provide 'writer_class_name', 'Traject::SolrPool::SolrJsonWriter'
provide 'solr.url', 'http://localhost:8983/solr/my_core'
provide 'solr_writer.thread_pool', 4
provide 'solr_pool.pool_size', 5
end
The origin (scheme://host:port) derived from solr.url becomes the pool's
base_url; the remainder of the URL becomes the relative request path. Basic
auth credentials are sent as an Authorization header, never baked into the
origin, so distinct credentials resolve to distinct pools and never share
connections.
Health check
writer = Traject::SolrPool::SolrJsonWriter.new(settings)
abort 'Solr unreachable' unless writer.ready?
ready? returns true for any HTTP response (the server answered — including
405/404), and false only on a transport failure (refused/timeout/DNS). It
uses a HEAD request so the pooled persistent connection stays clean for reuse.
Use writer.ping for the raw response. Configure the endpoint with
solr_writer.ping_path (default <core>/admin/ping) and the timeout with
solr_writer.ping_timeout (default 5s). See docs/usage.md for
the full reference, including reusing the pool from another class.
Settings
The writer honours the stock solr.* / solr_writer.* vocabulary, plus one
new namespaced setting.
| Setting | Default | Role |
|---|---|---|
solr.url |
(required unless solr.update_url set) |
Base Solr URL; /update/json is derived from it |
solr.update_url |
derived from solr.url |
Full update handler URL; used verbatim when provided |
solr_writer.batch_size |
100 |
Documents per batched update |
solr_writer.thread_pool |
1 |
Number of background writer threads |
solr_writer.max_skipped |
0 |
Skip tolerance before MaxSkippedRecordsExceeded (negative disables the cap) |
solr_writer.skippable_exceptions |
http.rb-native list (see below) | Exceptions treated as skippable during per-record retry |
solr_writer.commit_on_close |
false |
Send a commit when the writer closes (legacy solrj_writer.commit_on_close also honoured) |
solr_writer.solr_update_args |
none | Query params (e.g. { 'commitWithin' => 1000 }) applied to every update and delete request |
solr_writer.commit_solr_update_args |
{ 'commit' => 'true' } |
Query params for the commit request |
solr_writer.commit_timeout |
600 (10 min) |
Read timeout applied to the commit request only, so a slow commit is not cut off by a short http_timeout |
solr_writer.basic_auth_user |
embedded URI user | Basic-auth user (overrides credentials embedded in the URL) |
solr_writer.basic_auth_password |
embedded URI password | Basic-auth password (overrides credentials embedded in the URL) |
solr_writer.http_timeout |
none | Per-request HTTP timeout, passed to the pool connection |
solr_writer.pool_timeout |
pool default | Max time to wait for a connection checkout from the pool |
solr_pool.pool_size |
solr_writer.thread_pool + 1 |
Pool capacity, sized so no writer thread starves on checkout while the caller thread can still flush/commit |
Dropped settings
Two solr_json_writer.* settings from the stock writer are HTTPClient-specific
and are not supported:
solr_json_writer.http_client— the pool owns connection construction, so there is noHTTPClientinstance to inject.solr_json_writer.use_packaged_certs— anHTTPClientssl_config quirk; http.rb uses the operating system's certificate store instead.
Error handling
-
Traject::SolrPool::SolrJsonWriter::BadHttpResponse(aRuntimeError) is raised on any non-200 response. Its#responseaccessor returns the raw http.rb response (use#code/#to_s), and it extracts Solr's JSONerror.msginto the message when present. -
The default
solr_writer.skippable_exceptionslist is http.rb-native:[ HTTP::TimeoutError, HttpConnectionPool::TimeoutError, HTTP::ConnectionError, SocketError, Errno::ECONNREFUSED, Traject::SolrPool::SolrJsonWriter::BadHttpResponse ]A batch that fails is retried record by record; a record that raises a skippable exception is logged and counted, aborting the run only once
solr_writer.max_skippedis exceeded. -
Setting
solr_writer.skippable_exceptionsexplicitly replaces the default list entirely with your own. -
commit,delete, anddelete_all!raise raw on failure; they are not part of the skip path. -
Credentials never appear in logs or error messages: auth lives in a header and logs show the origin only.
Pool lifecycle
close flushes any queued records, waits for the background threads, and
commits when solr_writer.commit_on_close is set. It then leaves the pool
warm in the http_connection_pool registry so later writers (or a future
reader) on the same origin reuse it.
Teardown is the host application's concern. Close the pools explicitly on process shutdown or after a fork, for example:
HttpConnectionPool::Registry.instance.close_all
Development
Run the full CI pipeline (offline bundler-audit, RuboCop, RSpec):
bundle exec rake ci
Individual tasks (rake spec, rake rubocop, rake bundle:audit:check) are
also available.
Contributing
Bug reports and pull requests are welcome. Releasing and pushing are maintainer responsibilities — contributors should not push tags or publish gems.
Releasing
Releases are published to RubyGems via OIDC Trusted Publishing on a version tag — no API key is stored. The maintainer:
- Bumps the version:
rake bump:patch(or:minor/:major). This editslib/traject/solr_pool/version.rband prints the next commands. - Commits and tags:
git commit -am 'Release vX.Y.Z' && git tag vX.Y.Z. - Pushes the tag:
git push && git push --tags.
Pushing the v*.*.* tag triggers .github/workflows/release.yml, which
verifies the tag matches Traject::SolrPool::VERSION, re-runs the specs,
publishes the gem, and creates a GitHub Release with the gem and its SHA-256 /
SHA-512 checksums attached. gem push and git push are maintainer actions;
CI never pushes on its own.
License
The gem is available as open source under the terms of the MIT License.