ghostcrawl — Ruby SDK
The official Ruby client for the GhostCrawl API. Collect web data at scale — scrape, crawl, search, and extract structured data.
Install
gem install ghostcrawl
Or in your Gemfile:
gem "ghostcrawl", "~> 2.1"
Requires Ruby 3.0+. No runtime dependencies — uses stdlib net/http only.
Quickstart
require "ghostcrawl"
# Token from constructor or GHOSTCRAWL_API_KEY env var
client = Ghostcrawl::Client.new(token: "gck_live_YOUR_KEY")
# Scrape a URL
result = client.scrape(url: "https://example.com", format: "markdown")
puts result["content"]
# Start a crawl run and wait for it to finish — the server blocks until the run
# is terminal (no client-side poll-with-sleep loop).
run = client.crawl_runs.start(url: "https://example.com", max_depth: 2, max_pages: 50,
wait: true, timeout: 300)
puts "#{run["run_id"]} → #{run["status"]}"
# Web search
results = client.search(query: "latest AI research", engine: "google", limit: 10)
puts results["results"].length
Authentication
# Option 1: pass token directly
client = Ghostcrawl::Client.new(token: "gck_live_YOUR_KEY")
# Option 2: set environment variable (recommended for production)
# export GHOSTCRAWL_API_KEY=gck_live_YOUR_KEY
client = Ghostcrawl::Client.new
Every request sends Authorization: Bearer <token>. This is the only auth scheme the API accepts.
Extract structured data
require "ghostcrawl"
client = Ghostcrawl::Client.new(token: "gck_live_YOUR_KEY")
data = client.extract(
url: "https://example.com/product",
schema: {
type: "object",
properties: {
name: { type: "string" },
price: { type: "number" },
description: { type: "string" }
}
}
)
puts "#{data["name"]} — $#{data["price"]}"
Crawl runs — wait for completion
Crawl runs are asynchronous. Instead of hand-writing a poll-with-sleep loop, use
the event-driven wait API: the server blocks until the run reaches a terminal
state (completed / failed / cancelled) and returns the finished run.
# Start and wait in one call — sends wait_until: "completed"
run = client.crawl_runs.start(
url: "https://example.com",
max_depth: 2,
max_pages: 50,
wait: true, # block until terminal (no client sleep loop)
timeout: 300 # give up after 300s, returning the still-running run
)
puts run["status"] # "completed" | "failed" | "cancelled"
# Or wait on a run you already started (e.g. from a webhook or another process):
run = client.crawl_runs.wait_for_completion(run_id, timeout: 300)
# Deep crawl supports the same wait: keyword:
run = client.crawl(url: "https://example.com", wait: true, timeout: 300)
Under the hood wait_for_completion chains server-blocking
GET /v1/crawl-runs/{run_id}?wait=true&timeout_s=N requests until the run is
terminal or your timeout deadline — there is no sleep between checks. On
timeout the current non-terminal run is returned, so you can simply call
wait_for_completion again to keep waiting.
Prefer webhooks for fire-and-forget crawls. Register an endpoint with
client.webhooks.create(url:, event_types: ["crawl.completed"]) and let GhostCrawl
POST the result to you instead of holding a connection open.
Error handling
require "ghostcrawl"
client = Ghostcrawl::Client.new(token: "gck_live_YOUR_KEY")
begin
result = client.scrape(url: "https://example.com")
rescue Ghostcrawl::AuthenticationError
puts "Invalid API key — check your token"
rescue Ghostcrawl::PaymentRequiredError
puts "Usage limit reached — upgrade your plan"
rescue Ghostcrawl::RateLimitError
puts "Rate limit reached — retry after a short delay"
rescue Ghostcrawl::APIError => e
puts "Server error: #{e.status_code}"
rescue Ghostcrawl::GhostcrawlError => e
puts "API error: #{e.} (#{e.status_code})"
end
Self-hosted
client = Ghostcrawl::Client.new(
token: "gck_live_YOUR_KEY",
base_url: "http://localhost:8080"
)
All resources
| Resource | Method / accessor | Key operations |
|---|---|---|
| Scraping | client.scrape(url:) |
Render and return page content |
| Web search | client.search(query:) |
Search Google, Bing, DuckDuckGo |
| Data extraction | client.extract(url:, schema:) |
Structured JSON from any page |
| Deep crawl | client.crawl(url:) |
Crawl a site depth-first |
| URL map | client.map(url:) |
Discover all reachable URLs |
| Account | client.me |
Get account info and usage |
| Crawl runs | client.crawl_runs |
start (with wait:), wait_for_completion, list, get, cancel |
| Sessions | client.sessions |
create, list, extend, release |
| Profiles | client.profiles |
list, get_profile, create, update, delete |
| Webhooks | client.webhooks |
list, get_webhook, create, delete, rotate_secret |
| Schedules | client.schedules |
list, get_schedule, create, delete |
| Datasets | client.datasets |
list, get_dataset, create, delete, rows, append |
| Recordings | client.recordings |
list, get_recording, delete |
| Key-Value Store | client.kv |
get, set, delete |
License
Proprietary — GhostCrawl Software License. See LICENSE.