Class: ContextDev::Models::BatchSubmitParams::Input::Crawl::Data::HTML::Options::Pdf
- Inherits:
-
Internal::Type::BaseModel
- Object
- Internal::Type::BaseModel
- ContextDev::Models::BatchSubmitParams::Input::Crawl::Data::HTML::Options::Pdf
- Defined in:
- lib/context_dev/models/batch_submit_params.rb,
sig/context_dev/models/batch_submit_params.rbs
Overview
Defined Under Namespace
Modules: Ocr, ShouldParse
Instance Attribute Summary collapse
-
#end_ ⇒ Integer?
Last 1-based PDF page to parse.
-
#ocr ⇒ Boolean, ...
When true, OCR the selected PDF pages that have no usable text layer (scans), replacing each recovered page's text with the OCR result while pages with a real text layer keep it.
-
#should_parse ⇒ Boolean, ...
When true, PDF URLs are fetched and parsed.
-
#start ⇒ Integer?
First 1-based PDF page to parse.
Instance Method Summary collapse
-
#initialize(end_: nil, ocr: nil, should_parse: nil, start: nil) ⇒ Object
constructor
Some parameter documentations has been truncated, see Pdf for more details.
- #to_hash ⇒ {
Methods inherited from Internal::Type::BaseModel
==, #==, #[], coerce, #deconstruct_keys, #deep_to_h, dump, fields, hash, #hash, inherited, inspect, #inspect, known_fields, optional, recursively_to_h, required, #to_h, #to_json, #to_s, to_sorbet_type, #to_yaml
Methods included from Internal::Type::Converter
#coerce, coerce, #dump, dump, #inspect, inspect, meta_info, new_coerce_state, type_info
Methods included from Internal::Util::SorbetRuntimeSupport
#const_missing, #define_sorbet_constant!, #sorbet_constant_defined?, #to_sorbet_type, to_sorbet_type
Constructor Details
#initialize(end_: nil, ocr: nil, should_parse: nil, start: nil) ⇒ Object
Some parameter documentations has been truncated, see ContextDev::Models::BatchSubmitParams::Input::Crawl::Data::HTML::Options::Pdf for more details.
PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.
|
|
# File 'lib/context_dev/models/batch_submit_params.rb', line 2285
|
Instance Attribute Details
#end_ ⇒ Integer?
Last 1-based PDF page to parse. When omitted, parsing ends at the last page. Must be greater than or equal to start when both are provided.
2257 |
# File 'lib/context_dev/models/batch_submit_params.rb', line 2257 optional :end_, Integer, api_name: :end |
#ocr ⇒ Boolean, ...
When true, OCR the selected PDF pages that have no usable text layer (scans), replacing each recovered page's text with the OCR result while pages with a real text layer keep it. Billed at 1 credit per page OCR actually recovered, on top of the base request cost. When false, no OCR runs.
2266 |
# File 'lib/context_dev/models/batch_submit_params.rb', line 2266 optional :ocr, union: -> { ContextDev::BatchSubmitParams::Input::Crawl::Data::HTML::Options::Pdf::Ocr } |
#should_parse ⇒ Boolean, ...
When true, PDF URLs are fetched and parsed. When false, PDF URLs are skipped and a 400 PDF_SKIPPED is returned.
2273 2274 2275 2276 2277 |
# File 'lib/context_dev/models/batch_submit_params.rb', line 2273 optional :should_parse, union: -> { ContextDev::BatchSubmitParams::Input::Crawl::Data::HTML::Options::Pdf::ShouldParse }, api_name: :shouldParse |
#start ⇒ Integer?
First 1-based PDF page to parse. When omitted, parsing starts at the first page.
2283 |
# File 'lib/context_dev/models/batch_submit_params.rb', line 2283 optional :start, Integer |
Instance Method Details
#to_hash ⇒ {
2753 |
# File 'sig/context_dev/models/batch_submit_params.rbs', line 2753
def to_hash: -> {
|