Class: ContextDev::Models::WebWebScrapeHTMLParams::Pdf

Inherits:
Internal::Type::BaseModel show all
Defined in:
lib/context_dev/models/web_web_scrape_html_params.rb,
sig/context_dev/models/web_web_scrape_html_params.rbs

Defined Under Namespace

Modules: Ocr, ShouldParse

Instance Attribute Summary collapse

Instance Method Summary collapse

Methods inherited from Internal::Type::BaseModel

==, #==, #[], coerce, #deconstruct_keys, #deep_to_h, dump, fields, hash, #hash, inherited, inspect, #inspect, known_fields, optional, recursively_to_h, required, #to_h, #to_json, #to_s, to_sorbet_type, #to_yaml

Methods included from Internal::Type::Converter

#coerce, coerce, #dump, dump, #inspect, inspect, meta_info, new_coerce_state, type_info

Methods included from Internal::Util::SorbetRuntimeSupport

#const_missing, #define_sorbet_constant!, #sorbet_constant_defined?, #to_sorbet_type, to_sorbet_type

Constructor Details

#initialize(end_: nil, ocr: nil, should_parse: nil, start: nil) ⇒ Object

Some parameter documentations has been truncated, see ContextDev::Models::WebWebScrapeHTMLParams::Pdf for more details.

PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.

Parameters:

  • end_ (Integer) (defaults to: nil)

    Last 1-based PDF page to parse. When omitted, parsing ends at the last page. Mus

  • ocr (Boolean, Symbol, ContextDev::Models::WebWebScrapeHTMLParams::Pdf::Ocr) (defaults to: nil)

    When true, detect and OCR images embedded in the selected PDF pages, inserting r

  • should_parse (Boolean, Symbol, ContextDev::Models::WebWebScrapeHTMLParams::Pdf::ShouldParse) (defaults to: nil)

    When true, PDF URLs are fetched and parsed. When false, PDF URLs are skipped and

  • start (Integer) (defaults to: nil)

    First 1-based PDF page to parse. When omitted, parsing starts at the first page.



# File 'lib/context_dev/models/web_web_scrape_html_params.rb', line 486

Instance Attribute Details

#end_Integer?

Last 1-based PDF page to parse. When omitted, parsing ends at the last page. Must be greater than or equal to start when both are provided.

Parameters:

  • (Integer)

Returns:

  • (Integer, nil)


461
# File 'lib/context_dev/models/web_web_scrape_html_params.rb', line 461

optional :end_, Integer, api_name: :end

#ocrBoolean, ...

When true, detect and OCR images embedded in the selected PDF pages, inserting recognized text at each image's position in page reading order while preserving the PDF text layer. This is separate from automatic scanned-PDF OCR fallback.

Returns:



469
# File 'lib/context_dev/models/web_web_scrape_html_params.rb', line 469

optional :ocr, union: -> { ContextDev::WebWebScrapeHTMLParams::Pdf::Ocr }

#should_parseBoolean, ...

When true, PDF URLs are fetched and parsed. When false, PDF URLs are skipped and a 400 WEBSITE_ACCESS_ERROR is returned.



476
477
478
# File 'lib/context_dev/models/web_web_scrape_html_params.rb', line 476

optional :should_parse,
union: -> { ContextDev::WebWebScrapeHTMLParams::Pdf::ShouldParse },
api_name: :shouldParse

#startInteger?

First 1-based PDF page to parse. When omitted, parsing starts at the first page.

Parameters:

  • (Integer)

Returns:

  • (Integer, nil)


484
# File 'lib/context_dev/models/web_web_scrape_html_params.rb', line 484

optional :start, Integer

Instance Method Details

#to_hash{

Returns:

  • ({)


623
# File 'sig/context_dev/models/web_web_scrape_html_params.rbs', line 623

def to_hash: -> {