Class: ContextDev::Models::WebWebScrapeMdParams::Pdf
- Inherits:
-
Internal::Type::BaseModel
- Object
- Internal::Type::BaseModel
- ContextDev::Models::WebWebScrapeMdParams::Pdf
- Defined in:
- lib/context_dev/models/web_web_scrape_md_params.rb,
sig/context_dev/models/web_web_scrape_md_params.rbs
Defined Under Namespace
Modules: Ocr, ShouldParse
Instance Attribute Summary collapse
-
#end_ ⇒ Integer?
Last 1-based PDF page to parse.
-
#ocr ⇒ Boolean, ...
When true, detect and OCR images embedded in the selected PDF pages, inserting recognized text at each image's position in page reading order while preserving the PDF text layer.
-
#should_parse ⇒ Boolean, ...
When true, PDF URLs are fetched and parsed.
-
#start ⇒ Integer?
First 1-based PDF page to parse.
Instance Method Summary collapse
-
#initialize(end_: nil, ocr: nil, should_parse: nil, start: nil) ⇒ Object
constructor
Some parameter documentations has been truncated, see Pdf for more details.
- #to_hash ⇒ {
Methods inherited from Internal::Type::BaseModel
==, #==, #[], coerce, #deconstruct_keys, #deep_to_h, dump, fields, hash, #hash, inherited, inspect, #inspect, known_fields, optional, recursively_to_h, required, #to_h, #to_json, #to_s, to_sorbet_type, #to_yaml
Methods included from Internal::Type::Converter
#coerce, coerce, #dump, dump, #inspect, inspect, meta_info, new_coerce_state, type_info
Methods included from Internal::Util::SorbetRuntimeSupport
#const_missing, #define_sorbet_constant!, #sorbet_constant_defined?, #to_sorbet_type, to_sorbet_type
Constructor Details
#initialize(end_: nil, ocr: nil, should_parse: nil, start: nil) ⇒ Object
Some parameter documentations has been truncated, see ContextDev::Models::WebWebScrapeMdParams::Pdf for more details.
PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.
|
|
# File 'lib/context_dev/models/web_web_scrape_md_params.rb', line 561
|
Instance Attribute Details
#end_ ⇒ Integer?
Last 1-based PDF page to parse. When omitted, parsing ends at the last page. Must be greater than or equal to start when both are provided.
536 |
# File 'lib/context_dev/models/web_web_scrape_md_params.rb', line 536 optional :end_, Integer, api_name: :end |
#ocr ⇒ Boolean, ...
When true, detect and OCR images embedded in the selected PDF pages, inserting recognized text at each image's position in page reading order while preserving the PDF text layer. This is separate from automatic scanned-PDF OCR fallback.
544 |
# File 'lib/context_dev/models/web_web_scrape_md_params.rb', line 544 optional :ocr, union: -> { ContextDev::WebWebScrapeMdParams::Pdf::Ocr } |
#should_parse ⇒ Boolean, ...
When true, PDF URLs are fetched and parsed. When false, PDF URLs are skipped and a 400 WEBSITE_ACCESS_ERROR is returned.
551 552 553 |
# File 'lib/context_dev/models/web_web_scrape_md_params.rb', line 551 optional :should_parse, union: -> { ContextDev::WebWebScrapeMdParams::Pdf::ShouldParse }, api_name: :shouldParse |
#start ⇒ Integer?
First 1-based PDF page to parse. When omitted, parsing starts at the first page.
559 |
# File 'lib/context_dev/models/web_web_scrape_md_params.rb', line 559 optional :start, Integer |
Instance Method Details
#to_hash ⇒ {
672 |
# File 'sig/context_dev/models/web_web_scrape_md_params.rbs', line 672
def to_hash: -> {
|