Class: Ghostcrawl::Client
- Inherits:
-
Object
- Object
- Ghostcrawl::Client
- Defined in:
- lib/ghostcrawl/client.rb
Overview
GhostCrawl idiomatic API client.
Delegates all HTTP transport, URL routing, serialization, and auth to the Kiota-generated canonical core (_generated/). This facade is the shipped API.
Constant Summary collapse
- DEFAULT_TIMEOUT =
Default read timeout (seconds). Browser-rendered scrapes/crawls are slow, and the underlying net/http default (60s) is too short for them.
300
Instance Attribute Summary collapse
- #crawl_runs ⇒ CrawlRunsClient readonly
- #datasets ⇒ DatasetsClient readonly
- #kv ⇒ KVClient readonly
- #profiles ⇒ ProfilesClient readonly
- #recordings ⇒ RecordingsClient readonly
- #schedules ⇒ SchedulesClient readonly
- #sessions ⇒ SessionsClient readonly
- #webhooks ⇒ WebhooksClient readonly
Instance Method Summary collapse
-
#agent(url: nil, instruction: nil, task: nil, **opts) ⇒ Hash
Execute an agent task via POST /v1/agent.
-
#content(url:, engine: "auto", **opts) ⇒ Hash
Render a URL and return the rendered-content JSON envelope as a Hash.
-
#crawl(url:, max_depth: 2, max_pages: 100, wait: false, timeout: 300, raise_on_result_error: true, **opts) ⇒ Hash
Start a deep crawl from a seed URL.
-
#extract(url:, schema:, raise_on_result_error: true, **opts) ⇒ Hash
Extract structured data from a URL using a JSON Schema.
-
#identity(claim_os:, claim_browser:, device_model: nil, viewport: nil, **opts) ⇒ Hash
Create a new identity envelope.
-
#initialize(token: nil, base_url: nil, timeout: nil) ⇒ Client
constructor
A new instance of Client.
-
#map(url:, **opts) ⇒ Hash
Map all URLs reachable from a seed URL.
-
#me ⇒ Hash
Get the current account's profile.
-
#pdf(url:, paper_format: "a4", landscape: false, engine: "auto", **opts) ⇒ String
Render a URL to a PDF document and return the raw
application/pdfbytes. -
#scrape(url:, format: "markdown", engine: "auto", javascript: true, extract_schema: nil, raise_on_result_error: true, **opts) ⇒ Hash
Scrape a single URL and return the rendered content.
-
#screenshot(url:, format: "png", full_page: false, screenshot_selector: nil, engine: "auto", **opts) ⇒ String
Capture a screenshot of a URL and return the raw
image/pngbytes. -
#search(query:, engine: "google", limit: 10, provider_key: nil, **opts) ⇒ Hash
Search the web and return results.
-
#storage_states ⇒ Hash
List the account's persisted session storage-states.
-
#usage ⇒ Hash
Get the current account's cost/usage report.
Constructor Details
#initialize(token: nil, base_url: nil, timeout: nil) ⇒ Client
Returns a new instance of Client.
957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 |
# File 'lib/ghostcrawl/client.rb', line 957 def initialize(token: nil, base_url: nil, timeout: nil) resolved_token = token || ENV.fetch("GHOSTCRAWL_API_KEY", nil) if resolved_token.nil? || resolved_token.empty? raise ArgumentError, "token is required — pass token: or set GHOSTCRAWL_API_KEY. " \ "Get your key at https://ghostcrawl.io" end resolved_base = (base_url || ENV.fetch("GHOSTCRAWL_BASE_URL", nil) || DEFAULT_BASE_URL).gsub(%r{/+$}, "") resolved_timeout = (timeout || ENV.fetch("GHOSTCRAWL_TIMEOUT", nil) || DEFAULT_TIMEOUT).to_i # Build the Kiota core via BaseBearerTokenAuthenticationProvider + FaradayRequestAdapter. # All HTTP, auth, serialization, and URL routing delegate to the generated core. auth_provider = MicrosoftKiotaAbstractions::BaseBearerTokenAuthenticationProvider.new( StaticTokenProvider.new(resolved_token) ) adapter = MicrosoftKiotaFaraday::FaradayRequestAdapter.new(auth_provider) adapter.set_base_url(resolved_base) # The default Faraday connection sets no timeout, so it inherits net/http's # 60s read timeout — too short for browser-rendered work. Raise it. if adapter.client.respond_to?(:options) && resolved_timeout.positive? adapter.client..timeout = resolved_timeout adapter.client..open_timeout = 30 end @core = Ghostcrawl::GhostcrawlClient.new(adapter) @v1 = @core.v1 # Kept so #me can re-issue the GET /v1/me request with the raw-JSON # (Binary) response factory instead of the typed MeResponse model. @adapter = adapter # CrawlRunsClient needs the raw adapter for the event-driven long-poll wait # (the generated item builder can't attach ?wait/&timeout_s query params). @crawl_runs = CrawlRunsClient.new(@v1, @adapter) @sessions = SessionsClient.new(@v1) # Sub-clients with void-DELETE endpoints (204 No Content) also need the raw # adapter so ResponseHelper.void_request! can work around the broken generated # delete builders. See ResponseHelper.void_request! for the two Kiota defects. @profiles = ProfilesClient.new(@v1, @adapter) @webhooks = WebhooksClient.new(@v1, @adapter) @schedules = SchedulesClient.new(@v1, @adapter) @datasets = DatasetsClient.new(@v1) @recordings = RecordingsClient.new(@v1, @adapter) @kv = KVClient.new(@v1) end |
Instance Attribute Details
#crawl_runs ⇒ CrawlRunsClient (readonly)
1010 1011 1012 |
# File 'lib/ghostcrawl/client.rb', line 1010 def crawl_runs @crawl_runs end |
#datasets ⇒ DatasetsClient (readonly)
1020 1021 1022 |
# File 'lib/ghostcrawl/client.rb', line 1020 def datasets @datasets end |
#kv ⇒ KVClient (readonly)
1024 1025 1026 |
# File 'lib/ghostcrawl/client.rb', line 1024 def kv @kv end |
#profiles ⇒ ProfilesClient (readonly)
1014 1015 1016 |
# File 'lib/ghostcrawl/client.rb', line 1014 def profiles @profiles end |
#recordings ⇒ RecordingsClient (readonly)
1022 1023 1024 |
# File 'lib/ghostcrawl/client.rb', line 1022 def recordings @recordings end |
#schedules ⇒ SchedulesClient (readonly)
1018 1019 1020 |
# File 'lib/ghostcrawl/client.rb', line 1018 def schedules @schedules end |
#sessions ⇒ SessionsClient (readonly)
1012 1013 1014 |
# File 'lib/ghostcrawl/client.rb', line 1012 def sessions @sessions end |
#webhooks ⇒ WebhooksClient (readonly)
1016 1017 1018 |
# File 'lib/ghostcrawl/client.rb', line 1016 def webhooks @webhooks end |
Instance Method Details
#agent(url: nil, instruction: nil, task: nil, **opts) ⇒ Hash
Execute an agent task via POST /v1/agent.
The agent capability is gated per account; when it is not enabled the API
replies 404 not_found — this returns that problem+json body as a Hash
(carrying "detail") rather than raising, so callers can branch on
result.key?("detail"). The LLM step is BYO / not a hosted agent loop
(RESEARCH §2) — this method only proves the client serialize→POST→parse
plumbing round-trips against the real route.
/v1/agent has no generated request builder, so this hand-builds a
MicrosoftKiotaAbstractions::RequestInformation and routes it through the
SAME adapter (base URL + bearer auth + transport) as the modeled calls.
1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 |
# File 'lib/ghostcrawl/client.rb', line 1254 def agent(url: nil, instruction: nil, task: nil, **opts) payload = task || { "instruction" => instruction, "start_url" => url } data = { "task" => payload }.merge(opts.transform_keys(&:to_s)) request_info = MicrosoftKiotaAbstractions::RequestInformation.new request_info.http_method = :POST request_info.url_template = "{+baseurl}/v1/agent" request_info.headers.try_add("Accept", "application/json") request_info.set_stream_content(JSON.generate(data), "application/json") bytes = ResponseHelper.binary_request!(@adapter, request_info) parsed = JSON.parse(bytes) parsed.is_a?(Hash) ? parsed : { "detail" => parsed } rescue Ghostcrawl::GhostcrawlError => e # Agent is account-gated: a 404 not_found is the documented "capability # disabled" answer, not a transport failure. Surface the problem+json body # as a Hash (carrying "detail") rather than raising. Any other status is a # real error and is re-raised. raise unless e.status_code == 404 body = e.body.to_s brace = body.index("{") decoded = if brace begin JSON.parse(body[brace..]) rescue StandardError nil end end decoded.is_a?(Hash) ? decoded : { "detail" => e., "code" => (e.code || "not_found") } end |
#content(url:, engine: "auto", **opts) ⇒ Hash
Render a URL and return the rendered-content JSON envelope as a Hash. Delegates to POST /v1/content.
/v1/content responds with application/json (+url+, +status+,
+format+, +status_code+, +bytes+); this returns the decoded Hash. It uses
the same hand-built MicrosoftKiotaAbstractions::RequestInformation +
shared adapter dispatch as #pdf, parsing the JSON body at the end.
1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 |
# File 'lib/ghostcrawl/client.rb', line 1225 def content(url:, engine: "auto", **opts) data = { "url" => url, "engine" => engine }.merge(opts.transform_keys(&:to_s)) request_info = MicrosoftKiotaAbstractions::RequestInformation.new request_info.http_method = :POST request_info.url_template = "{+baseurl}/v1/content" request_info.headers.try_add("Accept", "application/json") request_info.set_stream_content(JSON.generate(data), "application/json") bytes = ResponseHelper.binary_request!(@adapter, request_info) parsed = JSON.parse(bytes) parsed.is_a?(Hash) ? parsed : { "content" => parsed } end |
#crawl(url:, max_depth: 2, max_pages: 100, wait: false, timeout: 300, raise_on_result_error: true, **opts) ⇒ Hash
Start a deep crawl from a seed URL. Delegates to POST /v1/crawl/deep via the generated CrawlDeepRequestBuilder.
Pass wait: true to block until the run is terminal. The deep-crawl start
returns a run_id, then the wait is handled by the same event-driven
server-blocking long-poll as Ghostcrawl::CrawlRunsClient#wait_for_completion — no
client-side sleep loop.
1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 |
# File 'lib/ghostcrawl/client.rb', line 1108 def crawl(url:, max_depth: 2, max_pages: 100, wait: false, timeout: 300, raise_on_result_error: true, **opts) data = { "seed_urls" => [url], "max_depth" => max_depth, "max_urls" => max_pages }.merge(opts.transform_keys(&:to_s)) hash = ResponseHelper.to_hash(@v1.crawl.deep.post(AdditionalDataBody.new(data))) unless wait return raise_on_result_error ? ResponseHelper.raise_on_result_error!(hash) : hash end run_id = hash["run_id"] return hash if run_id.nil? || run_id.to_s.empty? @crawl_runs.wait_for_completion(run_id, timeout: timeout) end |
#extract(url:, schema:, raise_on_result_error: true, **opts) ⇒ Hash
Extract structured data from a URL using a JSON Schema. Delegates to POST /v1/extract via the generated ExtractRequestBuilder.
1085 1086 1087 1088 1089 |
# File 'lib/ghostcrawl/client.rb', line 1085 def extract(url:, schema:, raise_on_result_error: true, **opts) data = { "url" => url, "schema" => schema }.merge(opts.transform_keys(&:to_s)) hash = ResponseHelper.to_hash(@v1.extract.post(AdditionalDataBody.new(data))) raise_on_result_error ? ResponseHelper.raise_on_result_error!(hash) : hash end |
#identity(claim_os:, claim_browser:, device_model: nil, viewport: nil, **opts) ⇒ Hash
Create a new identity envelope. Delegates to POST /v1/identity. Returns the opaque encrypted identity envelope; the SDK never decodes or inspects the ciphertext.
1293 1294 1295 1296 1297 1298 1299 |
# File 'lib/ghostcrawl/client.rb', line 1293 def identity(claim_os:, claim_browser:, device_model: nil, viewport: nil, **opts) data = { "claim_os" => claim_os, "claim_browser" => claim_browser } data["device_model"] = device_model unless device_model.nil? data["viewport"] = unless .nil? data.merge!(opts.transform_keys(&:to_s)) json_request(:POST, "/v1/identity", data) end |
#map(url:, **opts) ⇒ Hash
Map all URLs reachable from a seed URL. Delegates to POST /v1/map via the generated MapRequestBuilder.
1127 1128 1129 1130 |
# File 'lib/ghostcrawl/client.rb', line 1127 def map(url:, **opts) data = { "url" => url }.merge(opts.transform_keys(&:to_s)) ResponseHelper.to_hash(@v1.map.post(AdditionalDataBody.new(data))) end |
#me ⇒ Hash
Get the current account's profile. Delegates to GET /v1/me via the generated MeRequestBuilder.
The generated @v1.me.get deserializes into the typed MeResponse model,
whose created_at composed-type member fails under the pinned Kiota JSON
parser ("Error during deserialization"). To avoid that, we re-issue the
SAME request (reusing the builder's URL/auth/header wiring via
to_get_request_information) with the Binary response factory — the
exact raw-JSON path the scrape builder already uses (see KiotaParseNodeFix).
1142 1143 1144 1145 |
# File 'lib/ghostcrawl/client.rb', line 1142 def me request_info = @v1.me.to_get_request_information(nil) ResponseHelper.to_hash(@adapter.send_async(request_info, Ghostcrawl::V1::Binary, {})) end |
#pdf(url:, paper_format: "a4", landscape: false, engine: "auto", **opts) ⇒ String
Render a URL to a PDF document and return the raw application/pdf bytes.
Delegates to POST /v1/pdf.
/v1/pdf responds with binary, not a JSON envelope, so this returns the
bytes verbatim (an ASCII-8BIT String) — write them straight to a file:
data = client.pdf(url: "https://example.com")
File.binwrite("page.pdf", data)
PDF output is Chrome-only; a request that resolves to a Firefox or WebKit
identity is rejected with 400 pdf_engine_unsupported (a GhostcrawlError).
/v1/pdf has no generated request builder, so this hand-builds a
MicrosoftKiotaAbstractions::RequestInformation and routes it through the
SAME adapter (base URL + bearer auth + transport) as the modeled calls.
1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 |
# File 'lib/ghostcrawl/client.rb', line 1168 def pdf(url:, paper_format: "a4", landscape: false, engine: "auto", **opts) data = { "url" => url, "paper_format" => paper_format, "landscape" => landscape, "engine" => engine }.merge(opts.transform_keys(&:to_s)) request_info = MicrosoftKiotaAbstractions::RequestInformation.new request_info.http_method = :POST request_info.url_template = "{+baseurl}/v1/pdf" request_info.headers.try_add("Accept", "application/pdf") request_info.set_stream_content(JSON.generate(data), "application/json") ResponseHelper.binary_request!(@adapter, request_info) end |
#scrape(url:, format: "markdown", engine: "auto", javascript: true, extract_schema: nil, raise_on_result_error: true, **opts) ⇒ Hash
Scrape a single URL and return the rendered content. Delegates to POST /v1/scrape via the generated ScrapeRequestBuilder.
1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 |
# File 'lib/ghostcrawl/client.rb', line 1040 def scrape(url:, format: "markdown", engine: "auto", javascript: true, extract_schema: nil, raise_on_result_error: true, **opts) # Use AdditionalDataBody to send only the fields we specify — the generated # ScrapeRequest model would serialize typed defaults (nulls + empty enums) that # cause 422 validation errors on the server. data = { "url" => url, "format" => format, "engine" => engine, "javascript_enabled" => javascript }.merge(opts.transform_keys(&:to_s)) data["extract_schema"] = extract_schema unless extract_schema.nil? hash = ResponseHelper.to_hash(@v1.scrape.post(AdditionalDataBody.new(data))) normalize_scrape_content(hash) raise_on_result_error ? ResponseHelper.raise_on_result_error!(hash) : hash end |
#screenshot(url:, format: "png", full_page: false, screenshot_selector: nil, engine: "auto", **opts) ⇒ String
Capture a screenshot of a URL and return the raw image/png bytes.
Delegates to POST /v1/screenshot.
/v1/screenshot responds with a binary image, not a JSON envelope, so this
returns the bytes verbatim (an ASCII-8BIT String) — write them straight to a
file:
data = client.screenshot(url: "https://example.com")
File.binwrite("page.png", data)
/v1/screenshot has no generated request builder, so this hand-builds a
MicrosoftKiotaAbstractions::RequestInformation and routes it through the
SAME adapter (base URL + bearer auth + transport) as the modeled calls,
mirroring #pdf exactly (the raw-bytes dispatch idiom).
1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 |
# File 'lib/ghostcrawl/client.rb', line 1200 def screenshot(url:, format: "png", full_page: false, screenshot_selector: nil, engine: "auto", **opts) data = { "url" => url, "format" => format, "full_page" => full_page, "engine" => engine } data["screenshot_selector"] = screenshot_selector unless screenshot_selector.nil? data.merge!(opts.transform_keys(&:to_s)) request_info = MicrosoftKiotaAbstractions::RequestInformation.new request_info.http_method = :POST request_info.url_template = "{+baseurl}/v1/screenshot" request_info.headers.try_add("Accept", "image/png") request_info.set_stream_content(JSON.generate(data), "application/json") ResponseHelper.binary_request!(@adapter, request_info) end |
#search(query:, engine: "google", limit: 10, provider_key: nil, **opts) ⇒ Hash
Search the web and return results. Delegates to POST /v1/search via the generated SearchRequestBuilder.
/v1/search requires your own search-backend API key (BYO; GhostCrawl
charges no markup). Pass it as provider_key — it is sent as the
X-Provider-Authorization: Bearer <provider_key> header the backend
requires. Without it the API replies 401 search_backend_key_missing.
1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 |
# File 'lib/ghostcrawl/client.rb', line 1065 def search(query:, engine: "google", limit: 10, provider_key: nil, **opts) data = { "query" => query, "engine" => engine, "limit" => limit }.merge(opts.transform_keys(&:to_s)) config = nil unless provider_key.nil? config = MicrosoftKiotaAbstractions::RequestConfiguration.new headers = MicrosoftKiotaAbstractions::RequestHeaders.new headers.add("X-Provider-Authorization", "Bearer #{provider_key}") config.headers = headers end ResponseHelper.to_hash(@v1.search.post(AdditionalDataBody.new(data), config)) end |
#storage_states ⇒ Hash
List the account's persisted session storage-states. Delegates to GET /v1/storage-states.
1311 1312 1313 |
# File 'lib/ghostcrawl/client.rb', line 1311 def storage_states json_request(:GET, "/v1/storage-states", nil) end |
#usage ⇒ Hash
Get the current account's cost/usage report. Delegates to GET /v1/me/usage (the canonical account-scoped usage endpoint).
1304 1305 1306 |
# File 'lib/ghostcrawl/client.rb', line 1304 def usage json_request(:GET, "/v1/me/usage", nil) end |