Vision API — Ruby client
Official Ruby client for Vision API — send an image or a PDF, describe the fields you want in plain language, get structured JSON back with a confidence level on every value.
- Website — https://visionapi.io
- Documentation — https://docs.visionapi.io
- API keys — https://app.visionapi.io/dashboard/keys
- Preset catalog — https://visionapi.io/presets
- Playground — https://visionapi.io/playground
- Support — https://support.visionapi.io · https://visionapi.io/contact-us
Install
# Gemfile
gem "vision_api"
gem install vision_api
Ruby 3.0+. No runtime dependencies — net/http, json and openssl are all standard
library.
Quick start
require "vision_api"
vision = VisionAPI.new # reads ENV["VISION_API_KEY"]
res = vision.analyze(file: "invoice.pdf", preset: "invoice")
res["result"]["invoice_id"]["value"] # => "A-10422"
res["result"]["total"]["value"] # => 1284.5, or nil if the invoice has no total
res["credits_used"] # => 3
Requests are metered in credits, per image and per selected PDF page — see pricing for current rates. Failures cost nothing: the reservation is released in full on any non-2xx, so there is no compensating logic to write.
Server-side only. There is no publishable key and no test mode — an API key is a live spending credential. Keep it in the environment (or Rails credentials), never in a repository and never in anything a browser downloads.
Reading a result
Responses are plain hashes with the wire's keys, so everything you already know about hashes applies. Two rules explain almost every surprise:
1. Every scalar is wrapped. {"value" => …, "confidence" => "low"|"mid"|"high"}. Read
res["result"]["total"]["value"], not res["result"]["total"].
2. A preset response contains every field of that preset — including the ones the
document does not carry, which come back as {"value" => nil, "confidence" => "low"}. A
key being present does not mean a value was found.
Line-item arrays are the one shape worth looking at twice. The array itself is not wrapped; each cell inside each row is:
{
"invoice_id" => { "value" => "A-10422", "confidence" => "high" },
"carrier" => { "value" => nil, "confidence" => "low" },
"line_item" => [
{ "description" => { "value" => "Widget", "confidence" => "high" },
"quantity" => { "value" => 2, "confidence" => "high" },
"amount" => { "value" => 25.0, "confidence" => "mid" } }
]
}
VisionAPI::Result covers the common readings, so you rarely have to spell that out:
include VisionAPI::Result # or call them as VisionAPI::Result.unwrap(…)
unwrap(res["result"])
# => {"invoice_id" => "A-10422", "carrier" => nil, "line_item" => [{"description" => "Widget", …}]}
unwrap(res["result"], drop_null: true) # only what was actually found
value(res["result"], "total", 0) # 1284.5, or 0 when absent
rows(res["result"], "line_item") # [] when the invoice has no lines
present(res["result"]) # ["invoice_id", "total", "line_item"]
missing(res["result"]) # ["carrier", …]
below_confidence(res["result"], "high") # fields to route to a human
What you can send
Exactly one file source per call:
vision.analyze(file: "invoice.pdf", preset: "invoice") # a path
vision.analyze(file: Pathname("invoice.pdf"), preset: "invoice") # a Pathname
vision.analyze(file: File.open("invoice.pdf", "rb"), …) # an open binary IO
vision.analyze(file: ["scan.png", bytes], …) # bytes + a name
vision.analyze(file_url: "https://example.com/invoice.pdf", …) # a public URL
vision.analyze(file_base64: encoded, …) # "data:" prefix optional
JPEG, PNG, WebP, TIFF and PDF, up to 20 MB and 50 pages. The type is detected from magic bytes — the filename is ignored.
Options
| Keyword | Default | What it does |
|---|---|---|
preset: |
— | A catalog name, or "auto" to let the API classify the file first (free). |
schema: |
— | Custom fields, alone or on top of a preset. |
schema_name: |
— | A schema saved in your dashboard. Excludes preset: and schema:. |
pages: |
all | PDF page selection, e.g. "1-3,7". You pay for selected pages only. |
language_hint: |
auto | ISO 639-1 code, e.g. "es". |
detail: |
"standard" |
"high" renders pages at higher resolution. Same cost, slower. |
output: |
"json" |
"text" returns raw OCR text instead of fields. |
include_raw_text: |
false |
Adds full_text, the whole transcription, alongside result. |
min_confidence: |
"low" |
Fields below the level come back nil, with confidence preserved. |
Custom fields
A schema is a flat hash: each key is a field name, each value describes what to extract. It is compiled before any credit moves, so a bad schema costs nothing.
res = vision.analyze(
file: "invoice.pdf",
preset: "invoice",
schema: {
# Plain form — the string is the description, type defaults to string.
"machine_serial" => 'Serial number of the machine being invoiced, without the "SN:" prefix',
# Typed form.
"total_net" => { "type" => "number", "description" => "Total before tax" },
"signed_on" => { "type" => "date", "description" => "Date the contract was signed" },
# Reserved key: injects fields into every row of the preset's line-item array.
"line_item" => { "lot_number" => "The lot number printed on the line, if present" }
}
)
Field names must match ^[a-z][a-z0-9_]{0,63}$. Types are string (default), number,
boolean, date, array and object. A custom name that collides with a preset field is
a 422 schema_field_conflict — rename it, or use the preset's own field.
Descriptions are the prompt. "The invoice number exactly as printed, without the #"
extracts better than "invoice number". Say what to do when the value is missing or
ambiguous if it matters.
Reuse a combination by saving it:
vision.create_schema("our-invoices", preset: "invoice", schema: { "machine_serial" => "…" })
vision.analyze(file: "invoice.pdf", schema_name: "our-invoices")
Picking a preset
28 presets ship with the API. Fetch the catalog rather than hardcoding field names from memory — presets are versioned, and the catalog is the source of truth:
vision.presets.each { |p| puts "#{p['name']} (#{p['kind']}) — #{p['field_count']} fields" }
vision.preset("invoice")["fields"].map { |f| f["name"] }
Three ways to choose:
# 1. You know what it is.
vision.analyze(file: "receipt.jpg", preset: "receipt")
# 2. You don't, and you want the data anyway. Classification is free.
res = vision.analyze(file: "unknown.pdf", preset: "auto")
res["detection"]["preset"] # what ran
res["detection"]["fallback"] # true = "shape unknown", not a match
res["detection"]["alternatives"] # the rest of the ranking, best first
# 3. The *type* is the decision — routing a mixed inbox, or refusing to spend
# on a 40-page PDF until you know what it is. Far cheaper than extracting.
guess = vision.detect(file: "unknown.pdf")
vision.analyze(file: "unknown.pdf", preset: guess["recommended"]) unless guess["fallback"]
detect reads page 1 only, so an image and a 300-page PDF cost the same, and it is metered
in batches rather than per call: most calls report credits_used 0 and an occasional one
carries the charge. See pricing for the rate.
Questions instead of fields
Up to 5 questions about one file, priced exactly like an extraction. The questions themselves are free.
res = vision.ask(
file: "photo.jpg",
questions: ["Is there a dog in the image?", "How many people are visible?"]
)
res["answers"].each do |a|
case a["verdict"]
when "yes", "no" then handle(a["verdict"])
when "uncertain" then flag_for_review(a) # the image does not settle it — a real answer
when "n/a" then puts a["answer"] # it wasn't a yes/no question
end
end
Long jobs: async and webhooks
Synchronous requests are killed at 60 seconds with a 504 sync_timeout. Anything that
might run longer — a long PDF, detail: "high", a batch — belongs on the queue.
# Submit, then poll. wait_for_task handles the loop and the failure case.
task = vision.analyze_and_wait(
file: "contract-80-pages.pdf",
preset: "contract",
pages: "1-50",
poll_interval: 2,
max_wait: 900,
on_poll: ->(t) { Rails.logger.info(t["status"]) }
)
# Or submit and walk away — the result comes to you.
ref = vision.analyze_async(
file: "contract.pdf",
preset: "contract",
webhook_url: "https://yourapp.com/hooks/vision"
)
Results stay retrievable for 7 days; after that get_task raises ResultExpiredError
(metadata survives, the payload does not).
Verifying a delivery
Deliveries are signed. Verify over the raw bytes before parsing — a re-serialized body has different bytes and will not match.
class VisionHooksController < ApplicationController
skip_before_action :verify_authenticity_token
def create
event = VisionAPI::Webhook.verify(
request.raw_post,
request.headers["X-Vision-Signature"],
ENV.fetch("VISION_WEBHOOK_SECRET")
)
VisionResultJob.perform_later(event) # any 2xx is success — ack fast, work afterwards
head :accepted
rescue VisionAPI::WebhookSignatureError
head :bad_request # never parse an unverified body
end
end
VisionAPI::Webhook.verify rejects a bad signature, a malformed header and a timestamp
more than 5 minutes old, and accepts a delivery if any v1= part matches — which is
what makes a secret rotation seamless. Get the secret from
https://app.visionapi.io/dashboard/webhooks. Failed deliveries retry at +1 m, +5 m,
+15 m and +40 m, then stop.
Errors
Every failure raises a subclass of VisionAPI::APIError carrying the HTTP status, the
stable code, and whatever details the endpoint attached. Rescue the class you mean, or
switch on code — never on the message text, which is prose and changes.
begin
res = vision.analyze(file: "scan.pdf", preset: "invoice")
rescue VisionAPI::InsufficientCreditsError => e
alert_ops("needs #{e.required}, has #{e.available}") # never retried — it cannot succeed
rescue VisionAPI::SyncTimeoutError
task = vision.analyze_and_wait(file: "scan.pdf", preset: "invoice")
rescue VisionAPI::UnsupportedTypeError
quarantine("not an image or a PDF")
rescue VisionAPI::APIError => e
Rails.logger.error("vision #{e.code} (#{e.status}) request_id=#{e.request_id}")
end
| Class | HTTP | Codes |
|---|---|---|
InvalidRequestError |
400 | invalid_request |
AuthenticationError |
401 | invalid_api_key, unauthorized |
InsufficientCreditsError |
402 | insufficient_credits — with #required / #available |
PermissionDeniedError |
403 | forbidden, email_not_verified |
NotFoundError |
404 | task_not_found, schema_not_found |
ConflictError |
409 | conflict |
ResultExpiredError |
410 | result_expired |
PayloadTooLargeError |
413 | file_too_large, page_limit_exceeded |
UnsupportedTypeError |
415 | unsupported_type |
UnprocessableError |
422 | pdf_encrypted, invalid_page_selection, invalid_schema, schema_field_conflict, too_many_questions |
RateLimitError |
429 | rate_limited — with #retry_after |
TooManyTasksError |
429 | too_many_tasks — the per-plan async concurrency cap, with #max_tasks. Subclasses RateLimitError, but is not auto-retried: it clears when one of your tasks finishes |
InternalError |
500 | internal_error — with #request_id |
ProviderError |
502 | provider_error |
SyncTimeoutError |
504 | sync_timeout |
UsageError (bad arguments), ConnectionError / TimeoutError (the request never got a
response) and TaskFailedError / TaskTimeoutError come from the client itself.
Retries and idempotency
The client retries 429, 500, 502 and network failures — three attempts by default, with the
server's own Retry-After honored on 429 and exponential backoff with jitter elsewhere.
Input errors and insufficient_credits are never retried, because they cannot succeed.
Every billable POST is sent with a generated Idempotency-Key, so a retried upload replays
the first response instead of paying twice. Supply your own when the caller may retry — a
Sidekiq job that re-runs, a queue that redelivers — because a fresh process generates a
fresh key:
vision.analyze(file: path, preset: "invoice", idempotency_key: "invoice-#{invoice.id}")
Reusing a key with a different payload raises ConflictError, which is the mechanism
working: it means the key already stands for something else.
Configuration
vision = VisionAPI.new(
api_key: ENV["VISION_API_KEY"], # default: ENV["VISION_API_KEY"]
base_url: "https://api.visionapi.io", # default; override for a self-hosted deployment
timeout: 120, # per request, seconds
max_retries: 3,
auto_idempotency: true,
headers: { "X-Trace-Id" => trace_id } # sent on every request
)
Every method takes per-call idempotency_key: and timeout:.
Account and usage
credits = vision.credits
credits["balance"] # buckets are spent in order: subscription → rollover → pack → welcome
vision.each_request(limit: 100) do |record|
puts [record["created_at"], record["endpoint"], record["preset"], record["credits_used"]].join(" ")
end
# each_request without a block returns an Enumerator, so this fetches one page:
vision.each_request.first(10)
Usage history is metadata only — never the file, never the extracted values. Uploaded files are never retained: a synchronous request holds yours in memory for the length of the call, and an async request stages it only until the worker finishes with it.
Limits
Same for everyone:
| Limit | Value |
|---|---|
| Max file size | 20 MB |
| Max PDF pages per request | 50 |
| Sync request timeout | 60 s |
Per plan:
| Limit | Free | Starter | Growth | Pro | Scale |
|---|---|---|---|---|---|
| Requests per minute, per key | 10 | 60 | 120 | 300 | 600 |
| Burst capacity | 20 | 120 | 240 | 600 | 1,200 |
| Concurrent async tasks | 1 | 4 | 8 | 16 | 32 |
| Active API keys per account | 1 | 5 | 10 | 20 | 50 |
| Saved schemas | 3 | 10 | 25 | 100 | unlimited |
Max questions per ask |
5 | 5 | 5 | 10 | 10 |
The rate-limit bucket is per API key, not per account — splitting a workload across
keys splits the limit too. The concurrency cap is per account and does not split that way:
over it, an async submission answers 429 too_many_tasks and is charged nothing.
Higher limits on paid plans: https://visionapi.io/pricing.
Examples
Runnable scripts in examples/:
| File | What it shows |
|---|---|
analyze.rb |
The smallest useful call, and how to read the result |
custom_schema.rb |
Custom fields, line-item injection, saved schemas |
detect_then_analyze.rb |
Routing a mixed inbox before spending on extraction |
async_batch.rb |
A folder of long PDFs, queued with bounded concurrency |
webhook_server.rb |
A verified receiver, with no framework |
ask.rb |
Visual Q&A and the verdict field |
export VISION_API_KEY=sk_live_…
ruby examples/analyze.rb invoice.pdf
Development
bundle install
rake test # offline: a stub server stands in for the API, no key needed
rubocop
Contributing
Issues and pull requests are welcome at https://github.com/devrobotlabs/visionapi-ruby. For anything about the API itself — a preset, a limit, an error code — https://support.visionapi.io reaches the team faster.
License
MIT © Vision API