Vision API — Ruby client

Official Ruby client for Vision API — send an image or a PDF, describe the fields you want in plain language, get structured JSON back with a confidence level on every value.

Gem Version license


Install

# Gemfile
gem "vision_api"
gem install vision_api

Ruby 3.0+. No runtime dependencies — net/http, json and openssl are all standard library.

Quick start

require "vision_api"

vision = VisionAPI.new # reads ENV["VISION_API_KEY"]

res = vision.analyze(file: "invoice.pdf", preset: "invoice")

res["result"]["invoice_id"]["value"]  # => "A-10422"
res["result"]["total"]["value"]       # => 1284.5, or nil if the invoice has no total
res["credits_used"]                   # => 3

Requests are metered in credits, per image and per selected PDF page — see pricing for current rates. Failures cost nothing: the reservation is released in full on any non-2xx, so there is no compensating logic to write.

Server-side only. There is no publishable key and no test mode — an API key is a live spending credential. Keep it in the environment (or Rails credentials), never in a repository and never in anything a browser downloads.


Reading a result

Responses are plain hashes with the wire's keys, so everything you already know about hashes applies. Two rules explain almost every surprise:

1. Every scalar is wrapped. {"value" => …, "confidence" => "low"|"mid"|"high"}. Read res["result"]["total"]["value"], not res["result"]["total"].

2. A preset response contains every field of that preset — including the ones the document does not carry, which come back as {"value" => nil, "confidence" => "low"}. A key being present does not mean a value was found.

Line-item arrays are the one shape worth looking at twice. The array itself is not wrapped; each cell inside each row is:

{
  "invoice_id" => { "value" => "A-10422", "confidence" => "high" },
  "carrier"    => { "value" => nil,       "confidence" => "low" },
  "line_item"  => [
    { "description" => { "value" => "Widget", "confidence" => "high" },
      "quantity"    => { "value" => 2,        "confidence" => "high" },
      "amount"      => { "value" => 25.0,     "confidence" => "mid" } }
  ]
}

VisionAPI::Result covers the common readings, so you rarely have to spell that out:

include VisionAPI::Result  # or call them as VisionAPI::Result.unwrap(…)

unwrap(res["result"])
# => {"invoice_id" => "A-10422", "carrier" => nil, "line_item" => [{"description" => "Widget", …}]}

unwrap(res["result"], drop_null: true)   # only what was actually found
value(res["result"], "total", 0)         # 1284.5, or 0 when absent
rows(res["result"], "line_item")         # [] when the invoice has no lines
present(res["result"])                   # ["invoice_id", "total", "line_item"]
missing(res["result"])                   # ["carrier", …]
below_confidence(res["result"], "high")  # fields to route to a human

What you can send

Exactly one file source per call:

vision.analyze(file: "invoice.pdf", preset: "invoice")            # a path
vision.analyze(file: Pathname("invoice.pdf"), preset: "invoice")  # a Pathname
vision.analyze(file: File.open("invoice.pdf", "rb"), …)           # an open binary IO
vision.analyze(file: ["scan.png", bytes], …)                      # bytes + a name
vision.analyze(file_url: "https://example.com/invoice.pdf", …)    # a public URL
vision.analyze(file_base64: encoded, …)                           # "data:" prefix optional

JPEG, PNG, WebP, TIFF and PDF, up to 20 MB and 50 pages. The type is detected from magic bytes — the filename is ignored.

Options

Keyword Default What it does
preset: A catalog name, or "auto" to let the API classify the file first (free).
schema: Custom fields, alone or on top of a preset.
schema_name: A schema saved in your dashboard. Excludes preset: and schema:.
pages: all PDF page selection, e.g. "1-3,7". You pay for selected pages only.
language_hint: auto ISO 639-1 code, e.g. "es".
detail: "standard" "high" renders pages at higher resolution. Same cost, slower.
output: "json" "text" returns raw OCR text instead of fields.
include_raw_text: false Adds full_text, the whole transcription, alongside result.
min_confidence: "low" Fields below the level come back nil, with confidence preserved.

Custom fields

A schema is a flat hash: each key is a field name, each value describes what to extract. It is compiled before any credit moves, so a bad schema costs nothing.

res = vision.analyze(
  file: "invoice.pdf",
  preset: "invoice",
  schema: {
    # Plain form — the string is the description, type defaults to string.
    "machine_serial" => 'Serial number of the machine being invoiced, without the "SN:" prefix',

    # Typed form.
    "total_net" => { "type" => "number", "description" => "Total before tax" },
    "signed_on" => { "type" => "date", "description" => "Date the contract was signed" },

    # Reserved key: injects fields into every row of the preset's line-item array.
    "line_item" => { "lot_number" => "The lot number printed on the line, if present" }
  }
)

Field names must match ^[a-z][a-z0-9_]{0,63}$. Types are string (default), number, boolean, date, array and object. A custom name that collides with a preset field is a 422 schema_field_conflict — rename it, or use the preset's own field.

Descriptions are the prompt. "The invoice number exactly as printed, without the #" extracts better than "invoice number". Say what to do when the value is missing or ambiguous if it matters.

Reuse a combination by saving it:

vision.create_schema("our-invoices", preset: "invoice", schema: { "machine_serial" => "" })
vision.analyze(file: "invoice.pdf", schema_name: "our-invoices")

Picking a preset

28 presets ship with the API. Fetch the catalog rather than hardcoding field names from memory — presets are versioned, and the catalog is the source of truth:

vision.presets.each { |p| puts "#{p['name']} (#{p['kind']}) — #{p['field_count']} fields" }
vision.preset("invoice")["fields"].map { |f| f["name"] }

Three ways to choose:

# 1. You know what it is.
vision.analyze(file: "receipt.jpg", preset: "receipt")

# 2. You don't, and you want the data anyway. Classification is free.
res = vision.analyze(file: "unknown.pdf", preset: "auto")
res["detection"]["preset"]        # what ran
res["detection"]["fallback"]      # true = "shape unknown", not a match
res["detection"]["alternatives"]  # the rest of the ranking, best first

# 3. The *type* is the decision — routing a mixed inbox, or refusing to spend
#    on a 40-page PDF until you know what it is. Far cheaper than extracting.
guess = vision.detect(file: "unknown.pdf")
vision.analyze(file: "unknown.pdf", preset: guess["recommended"]) unless guess["fallback"]

detect reads page 1 only, so an image and a 300-page PDF cost the same, and it is metered in batches rather than per call: most calls report credits_used 0 and an occasional one carries the charge. See pricing for the rate.


Questions instead of fields

Up to 5 questions about one file, priced exactly like an extraction. The questions themselves are free.

res = vision.ask(
  file: "photo.jpg",
  questions: ["Is there a dog in the image?", "How many people are visible?"]
)

res["answers"].each do |a|
  case a["verdict"]
  when "yes", "no" then handle(a["verdict"])
  when "uncertain" then flag_for_review(a)  # the image does not settle it — a real answer
  when "n/a"       then puts a["answer"]    # it wasn't a yes/no question
  end
end

Long jobs: async and webhooks

Synchronous requests are killed at 60 seconds with a 504 sync_timeout. Anything that might run longer — a long PDF, detail: "high", a batch — belongs on the queue.

# Submit, then poll. wait_for_task handles the loop and the failure case.
task = vision.analyze_and_wait(
  file: "contract-80-pages.pdf",
  preset: "contract",
  pages: "1-50",
  poll_interval: 2,
  max_wait: 900,
  on_poll: ->(t) { Rails.logger.info(t["status"]) }
)

# Or submit and walk away — the result comes to you.
ref = vision.analyze_async(
  file: "contract.pdf",
  preset: "contract",
  webhook_url: "https://yourapp.com/hooks/vision"
)

Results stay retrievable for 7 days; after that get_task raises ResultExpiredError (metadata survives, the payload does not).

Verifying a delivery

Deliveries are signed. Verify over the raw bytes before parsing — a re-serialized body has different bytes and will not match.

class VisionHooksController < ApplicationController
  skip_before_action :verify_authenticity_token

  def create
    event = VisionAPI::Webhook.verify(
      request.raw_post,
      request.headers["X-Vision-Signature"],
      ENV.fetch("VISION_WEBHOOK_SECRET")
    )

    VisionResultJob.perform_later(event) # any 2xx is success — ack fast, work afterwards
    head :accepted
  rescue VisionAPI::WebhookSignatureError
    head :bad_request                    # never parse an unverified body
  end
end

VisionAPI::Webhook.verify rejects a bad signature, a malformed header and a timestamp more than 5 minutes old, and accepts a delivery if any v1= part matches — which is what makes a secret rotation seamless. Get the secret from https://app.visionapi.io/dashboard/webhooks. Failed deliveries retry at +1 m, +5 m, +15 m and +40 m, then stop.


Errors

Every failure raises a subclass of VisionAPI::APIError carrying the HTTP status, the stable code, and whatever details the endpoint attached. Rescue the class you mean, or switch on code — never on the message text, which is prose and changes.

begin
  res = vision.analyze(file: "scan.pdf", preset: "invoice")
rescue VisionAPI::InsufficientCreditsError => e
  alert_ops("needs #{e.required}, has #{e.available}")  # never retried — it cannot succeed
rescue VisionAPI::SyncTimeoutError
  task = vision.analyze_and_wait(file: "scan.pdf", preset: "invoice")
rescue VisionAPI::UnsupportedTypeError
  quarantine("not an image or a PDF")
rescue VisionAPI::APIError => e
  Rails.logger.error("vision #{e.code} (#{e.status}) request_id=#{e.request_id}")
end
Class HTTP Codes
InvalidRequestError 400 invalid_request
AuthenticationError 401 invalid_api_key, unauthorized
InsufficientCreditsError 402 insufficient_credits — with #required / #available
PermissionDeniedError 403 forbidden, email_not_verified
NotFoundError 404 task_not_found, schema_not_found
ConflictError 409 conflict
ResultExpiredError 410 result_expired
PayloadTooLargeError 413 file_too_large, page_limit_exceeded
UnsupportedTypeError 415 unsupported_type
UnprocessableError 422 pdf_encrypted, invalid_page_selection, invalid_schema, schema_field_conflict, too_many_questions
RateLimitError 429 rate_limited — with #retry_after
TooManyTasksError 429 too_many_tasks — the per-plan async concurrency cap, with #max_tasks. Subclasses RateLimitError, but is not auto-retried: it clears when one of your tasks finishes
InternalError 500 internal_error — with #request_id
ProviderError 502 provider_error
SyncTimeoutError 504 sync_timeout

UsageError (bad arguments), ConnectionError / TimeoutError (the request never got a response) and TaskFailedError / TaskTimeoutError come from the client itself.

Retries and idempotency

The client retries 429, 500, 502 and network failures — three attempts by default, with the server's own Retry-After honored on 429 and exponential backoff with jitter elsewhere. Input errors and insufficient_credits are never retried, because they cannot succeed.

Every billable POST is sent with a generated Idempotency-Key, so a retried upload replays the first response instead of paying twice. Supply your own when the caller may retry — a Sidekiq job that re-runs, a queue that redelivers — because a fresh process generates a fresh key:

vision.analyze(file: path, preset: "invoice", idempotency_key: "invoice-#{invoice.id}")

Reusing a key with a different payload raises ConflictError, which is the mechanism working: it means the key already stands for something else.


Configuration

vision = VisionAPI.new(
  api_key: ENV["VISION_API_KEY"],       # default: ENV["VISION_API_KEY"]
  base_url: "https://api.visionapi.io", # default; override for a self-hosted deployment
  timeout: 120,                          # per request, seconds
  max_retries: 3,
  auto_idempotency: true,
  headers: { "X-Trace-Id" => trace_id }  # sent on every request
)

Every method takes per-call idempotency_key: and timeout:.


Account and usage

credits = vision.credits
credits["balance"] # buckets are spent in order: subscription → rollover → pack → welcome

vision.each_request(limit: 100) do |record|
  puts [record["created_at"], record["endpoint"], record["preset"], record["credits_used"]].join(" ")
end

# each_request without a block returns an Enumerator, so this fetches one page:
vision.each_request.first(10)

Usage history is metadata only — never the file, never the extracted values. Uploaded files are never retained: a synchronous request holds yours in memory for the length of the call, and an async request stages it only until the worker finishes with it.


Limits

Same for everyone:

Limit Value
Max file size 20 MB
Max PDF pages per request 50
Sync request timeout 60 s

Per plan:

Limit Free Starter Growth Pro Scale
Requests per minute, per key 10 60 120 300 600
Burst capacity 20 120 240 600 1,200
Concurrent async tasks 1 4 8 16 32
Active API keys per account 1 5 10 20 50
Saved schemas 3 10 25 100 unlimited
Max questions per ask 5 5 5 10 10

The rate-limit bucket is per API key, not per account — splitting a workload across keys splits the limit too. The concurrency cap is per account and does not split that way: over it, an async submission answers 429 too_many_tasks and is charged nothing. Higher limits on paid plans: https://visionapi.io/pricing.


Examples

Runnable scripts in examples/:

File What it shows
analyze.rb The smallest useful call, and how to read the result
custom_schema.rb Custom fields, line-item injection, saved schemas
detect_then_analyze.rb Routing a mixed inbox before spending on extraction
async_batch.rb A folder of long PDFs, queued with bounded concurrency
webhook_server.rb A verified receiver, with no framework
ask.rb Visual Q&A and the verdict field
export VISION_API_KEY=sk_live_…
ruby examples/analyze.rb invoice.pdf

Development

bundle install
rake test        # offline: a stub server stands in for the API, no key needed
rubocop

Contributing

Issues and pull requests are welcome at https://github.com/devrobotlabs/visionapi-ruby. For anything about the API itself — a preset, a limit, an error code — https://support.visionapi.io reaches the team faster.

License

MIT © Vision API