Class: Wp2txt::Corpus
- Inherits:
-
Object
- Object
- Wp2txt::Corpus
- Defined in:
- lib/wp2txt/corpus.rb
Overview
Facade over a local dump: single-article access (Tier 0, multistream), exhaustive metadata queries (Tier 1, MetadataIndex), and corpus extraction. Shared by the CLI and the MCP server so both expose identical behavior.
Defined Under Namespace
Classes: Cancelled, LazyTitleIndex
Constant Summary collapse
- DEFAULT_MAX_SYNC_ARTICLES =
Sync extraction cap: larger requests need the (future) job API
5000- RENDER_CONFIG =
{ format: :text, title: true, heading: true, list: false, table: false, pre: false, ref: false, redirect: false, multiline: false, category: true, category_only: false, summary_only: false, metadata_only: false, marker: true, markers: true, extract_citations: false, expand_templates: true, sections: nil, section_output: "structured", min_section_length: 0, skip_empty: false, alias_file: nil, no_section_aliases: false, show_matched_sections: false }.freeze
- DEFAULT_MAX_CHARS =
Default cap on article text returned inline (LLM context economy); callers can raise it explicitly, and truncation is always flagged
40_000- GUARDRAIL_MAX_RATIO =
Guardrail thresholds: pairs whose co-occurrence ratio exceeds GUARDRAIL_MAX_RATIO (with both headings above GUARDRAIL_MIN_ARTICLES) coexist in the same articles and are likely NOT synonyms
0.2- GUARDRAIL_MIN_ARTICLES =
100- SNIPPET_CONTEXT =
Save an alias set after mechanically verifying each group: high co-occurrence pairs block the save unless force is set, so protocol compliance does not depend on the calling model's discipline.
Tier 2: full-text search
80- EXTRACT_BATCH_SIZE =
Titles fetched/rendered per batch while streaming to disk
200- SQL_ROW_LIMIT =
Read-only SQL (escape hatch for queries the fixed tools cannot express)
200- SQL_CELL_LIMIT =
2000- SQL_TIMEOUT_SECONDS =
30- SQL_FORBIDDEN =
/\b(ATTACH|DETACH|PRAGMA|INSERT|UPDATE|DELETE|DROP|CREATE|ALTER|REPLACE|VACUUM|REINDEX)\b/i
Constants included from Wp2txt
AUTHORITY_CONTROL_REGEX, AUTHORITY_CONTROL_TEMPLATES, BLANK_LINE_REGEX, BLOCK_EXTENSION_TAGS, BLOCK_MARKERS, BYTES_PER_GB, BYTES_PER_KB, BYTES_PER_MB, CATEGORY_LINE_REGEX, CATEGORY_NAMESPACES, CATEGORY_NAMESPACE_REGEX, CATEGORY_REGEX, CHRREF_TO_UTF_REGEX, CITATION_TEMPLATES, CITATION_TEMPLATE_REGEX, CLEANUP_EMPTY_PARENS_REGEX, CLEANUP_FILE_INCOMPLETE_REGEX, CLEANUP_FILE_LINE_REGEX, CLEANUP_KEY_VALUE_LINE_REGEX, CLEANUP_MIXED_WHITESPACE_REGEX, CLEANUP_MULTIPLE_PIPES_REGEX, CLEANUP_MULTIPLE_SPACES_REGEX, CLEANUP_MULTI_BLANK_REGEX, CLEANUP_ORPHANED_BRACKETS_REGEX, CLEANUP_ORPHANED_CLOSE_REGEX, CLEANUP_ORPHANED_OPEN_REGEX, CLEANUP_PIPE_CLOSE_REGEX, CLEANUP_PIPE_LINE_REGEX, CLEANUP_REGEX_01, CLEANUP_REGEX_02, CLEANUP_REGEX_03, CLEANUP_REGEX_04, CLEANUP_REGEX_05, CLEANUP_REGEX_06, CLEANUP_REGEX_07, CLEANUP_REGEX_08, CLEANUP_REMNANTS_REGEX, CLEANUP_REMNANTS_TEMPLATES, CLEANUP_STANDALONE_CLOSE_REGEX, CLEANUP_TRAILING_PIPE_REGEX, COMPLEX_REGEX_01, COMPLEX_REGEX_02, COMPLEX_REGEX_03, COMPLEX_REGEX_04, COMPLEX_REGEX_05, CONVERT_TEMPLATES, COUNTRY_CODE_REGEX, CURLY_SQUARE_BRACKET_REGEX, DEFAULTSORT_KEYWORDS, DEFAULT_BUFFER_SIZE, DEFAULT_CATEGORY_CACHE_EXPIRY_DAYS, DEFAULT_DUMP_EXPIRY_DAYS, DEFAULT_HTTP_TIMEOUT, DEFAULT_MARKERS, DEFAULT_PROGRESS_INTERVAL, DEFAULT_TOP_N_SECTIONS, DEF_MARKS_REGEX, DISPLAYTITLE_KEYWORDS, DOUBLE_CURLY_BRACKET_REGEX, DOUBLE_SQUARE_BRACKET_REGEX, DOUBLE_UNDERSCORE_MAGIC_REGEX, DOUBLE_UNDERSCORE_PATTERNS, ENTITIES, ESCAPE_NOWIKI_REGEX, EXTENSION_TAGS, EXTRA_ENTITIES, EXTRA_ENTITIES_REGEX, FILE_NAMESPACES, FILE_NAMESPACES_REGEX, FLAG_TEMPLATES, FLAG_TEMPLATE_REGEX, FORMATTING_TEMPLATES, FORMATTING_TEMPLATE_REGEX, FORMAT_REF_REGEX, HEADING_CODA_REGEX, HEADING_ONSET_REGEX, HTML_COMMENT_REGEX, HTML_DECODER, HTML_ENTITIES_PATH, HTML_HASH, HTML_REGEX, IMAGEMAP_COORD_REGEX, IMAGE_PARAMS, IMAGE_PARAMS_REGEX, IMAGE_PARAM_KEYS, INDEX_PROGRESS_THRESHOLD, INLINE_MARKERS, INTERWIKI_LINK_TEMPLATES, INTERWIKI_PREFIX_REGEX, IN_DEFINITION_REGEX, IN_HEADING_REGEX, IN_HTML_TABLE_REGEX, IN_HTML_TABLE_REGEX1, IN_HTML_TABLE_REGEX2, IN_INPUTBOX_REGEX, IN_INPUTBOX_REGEX1, IN_INPUTBOX_REGEX2, IN_LINK_REGEX, IN_MATH_REGEX, IN_MATH_REGEX1, IN_MATH_REGEX2, IN_ORDERED_REGEX, IN_PRE_REGEX, IN_SOURCE_REGEX, IN_SOURCE_REGEX1, IN_SOURCE_REGEX2, IN_TABLE_REGEX1, IN_TABLE_REGEX2, IN_UNORDERED_REGEX, ISOLATED_TAG_REGEX, ISOLATED_TEMPLATE_REGEX, LIST_MARKS_REGEX, LONE_ASTERISK_REGEX, MAGIC_WORD_LINE_REGEX, MAGIC_WORD_TEMPLATE_REGEX, MAKE_REFERENCE_REGEX_A, MAKE_REFERENCE_REGEX_B, MAKE_REFERENCE_REGEX_C, MAKE_REFERENCE_REGEX_D, MARKER_PATTERNS, MARKER_TYPES, MATH_ENTITIES, MATH_ENTITIES_REGEX, MAX_BUFFER_SIZE, MAX_HTTP_RETRIES, MAX_NESTING_ITERATIONS, MEDIAWIKI_DATA_PATH, MIN_BUFFER_SIZE, MIXED_SCRIPT_TEMPLATES, ML_LINK_END_REGEX, ML_LINK_ONSET_REGEX, ML_TEMPLATE_END_REGEX, ML_TEMPLATE_ONSET_REGEX, MNDASH_REGEX, NON_ARTICLE_NAMESPACES, NON_ARTICLE_NAMESPACE_REGEX, ONSET_BAR_REGEX, PRE_MARKS_REGEX, RACTOR_SHAREABLE_EXCLUDES, REDIRECT_KEYWORDS, REDIRECT_REGEX, REMOVE_DIRECTIVES_REGEX, REMOVE_DIRECTIVES_REGEX_GENERIC, REMOVE_EMPHASIS_REGEX, REMOVE_HR_REGEX, REMOVE_INLINE_REGEX, REMOVE_ISOLATED_REGEX, REMOVE_TAG_REGEX, REMOVE_TEMPLATES, REMOVE_TEMPLATES_REGEX, RESUME_METADATA_MAX_AGE_DAYS, RUBY_TEXT_TEMPLATES, SECONDS_PER_DAY, SECONDS_PER_HOUR, SECONDS_PER_MINUTE, SELF_CLOSING_TAG_REGEX, SINGLE_CURLY_BRACKET_REGEX, SINGLE_SQUARE_BRACKET_REGEX, SISTER_PROJECTS, TEMPLATE_DATA_PATH, UNESCAPE_NOWIKI_REGEX, VERSION, WIKIMEDIA_PROJECT_NAMES, WIKIMEDIA_PROJECT_REGEX, WIKIPEDIA_ENTITIES_PATH
Instance Attribute Summary collapse
-
#index_path ⇒ Object
readonly
Returns the value of attribute index_path.
-
#metadata ⇒ Object
readonly
Returns the value of attribute metadata.
-
#multistream_path ⇒ Object
readonly
Returns the value of attribute multistream_path.
Class Method Summary collapse
- .for_input(multistream_path, cache_dir: nil) ⇒ Object
-
.for_lang(lang, cache_dir: nil) ⇒ Object
Build a Corpus from a language code using the DumpManager cache.
Instance Method Summary collapse
- #category_tree(category, depth: 2) ⇒ Object
- #close ⇒ Object
-
#describe_schema ⇒ Object
CREATE statements of all tables available to query_sql.
- #dump_info ⇒ Object
- #extract_corpus(output_path:, content: "sections", sections: nil, alias_set: nil, category: nil, depth: 0, categories: nil, category_match: nil, title_match: nil, limit: 0, chunk_size: nil, chunk_overlap: 0, max_articles: DEFAULT_MAX_SYNC_ARTICLES, num_processes: 4, progress: nil, cancel_check: nil) ⇒ Object
-
#find_articles(**filters) ⇒ Object
------------------------------------------------------------------ Tier 1: exhaustive queries (delegated to MetadataIndex) ------------------------------------------------------------------.
- #fts ⇒ Object
- #get_alias_set(name) ⇒ Object
- #get_article(title, format: "text", follow_redirect: true, max_chars: DEFAULT_MAX_CHARS) ⇒ Object
-
#get_categories(title) ⇒ Object
Categories of one article, resolved through the same title normalization.
-
#get_sections(title, sections, alias_set: nil) ⇒ Object
Extract specific sections from one article.
-
#initialize(multistream_path:, index_path:, cache_dir: nil) ⇒ Corpus
constructor
A new instance of Corpus.
- #list_alias_sets ⇒ Object
- #list_headings(title) ⇒ Object
-
#metadata_built? ⇒ Boolean
------------------------------------------------------------------ Info ------------------------------------------------------------------.
-
#query_sql(sql, limit: SQL_ROW_LIMIT, timeout: SQL_TIMEOUT_SECONDS) ⇒ Object
Run a read-only SELECT against the metadata DB (with the FTS DB attached as
ftswhen built). - #save_alias_set(name, groups, force: false, max_ratio: GUARDRAIL_MAX_RATIO, min_articles: GUARDRAIL_MIN_ARTICLES) ⇒ Object
-
#search_text(query, mode: "phrase", sections: nil, alias_set: nil, category: nil, depth: 0, limit: 20, offset: 0, count: "capped", snippets: true) ⇒ Object
Exhaustive full-text search over cleaned section text.
- #section_cooccurrence(headings, category: nil, depth: 0) ⇒ Object
- #section_stats(category: nil, depth: 0, top_n: 50) ⇒ Object
Methods included from Formatter
#build_text_content, #extract_redirect, #format_article, #format_article_json, #format_article_text, #format_category_only, #format_full_article, #format_metadata_only, #format_metadata_only_json, #format_metadata_only_text, #format_sections_combined_json, #format_sections_combined_text, #format_sections_structured_json, #format_sections_structured_text, #format_with_categories, #format_with_sections, #formatter_debug_mode, #process_element
Methods included from Wp2txt
#apply_markers, #apply_pipe_trick, article_page?, #batch_file_mod, build_extra_entities_regex, build_template_regex, #chrref_to_utf, #cleanup, #collect_files, #convert_characters, #correct_inline_template, #correct_separator, days_to_seconds, #escape_nowiki, #extract_template_content, file_age_days, file_fresh?, #file_mod, #finalize_markers, #format_citation, format_file_size, #format_wiki, #format_wiki_regex_transform, #html_decoder, load_html_entities, load_mediawiki_data, load_template_data, make_constants_ractor_shareable!, #make_reference, #marker_placeholder, #mndash, #parse_markers_config, #process_external_links, #process_interwiki_links, #process_nested_single_pass, #process_nested_structure, #remove_complex, #remove_directive, #remove_emphasis, #remove_hr, #remove_html, #remove_inbetween, #remove_ref, #remove_table, #remove_tag, #remove_templates, #rename, #replace_paired_templates_with_marker, #replace_template_with_marker, #replace_wiki_table_with_marker, #sec_to_str, #special_chr, ssl_safe_get, #template_matches?, #unescape_nowiki
Constructor Details
#initialize(multistream_path:, index_path:, cache_dir: nil) ⇒ Corpus
Returns a new instance of Corpus.
60 61 62 63 64 65 |
# File 'lib/wp2txt/corpus.rb', line 60 def initialize(multistream_path:, index_path:, cache_dir: nil) @multistream_path = multistream_path @index_path = index_path @cache_dir = cache_dir @metadata = MetadataIndex.new(MetadataIndex.path_for(multistream_path, cache_dir: cache_dir)) end |
Instance Attribute Details
#index_path ⇒ Object (readonly)
Returns the value of attribute index_path.
58 59 60 |
# File 'lib/wp2txt/corpus.rb', line 58 def index_path @index_path end |
#metadata ⇒ Object (readonly)
Returns the value of attribute metadata.
58 59 60 |
# File 'lib/wp2txt/corpus.rb', line 58 def @metadata end |
#multistream_path ⇒ Object (readonly)
Returns the value of attribute multistream_path.
58 59 60 |
# File 'lib/wp2txt/corpus.rb', line 58 def multistream_path @multistream_path end |
Class Method Details
.for_input(multistream_path, cache_dir: nil) ⇒ Object
80 81 82 83 84 85 86 87 88 89 90 |
# File 'lib/wp2txt/corpus.rb', line 80 def self.for_input(multistream_path, cache_dir: nil) candidates = [ multistream_path.sub(/multistream\.xml\.bz2\z/, "multistream-index.txt.bz2"), multistream_path.sub(/\.xml\.bz2\z/, "-index.txt.bz2"), multistream_path.sub(/\.xml\.bz2\z/, "-index.txt") ].uniq index = candidates.find { |c| c != multistream_path && File.exist?(c) } raise ArgumentError, "Multistream index file not found next to #{multistream_path}" unless index new(multistream_path: multistream_path, index_path: index, cache_dir: cache_dir) end |
.for_lang(lang, cache_dir: nil) ⇒ Object
Build a Corpus from a language code using the DumpManager cache. Raises with guidance when the dump has not been downloaded yet.
69 70 71 72 73 74 75 76 77 78 |
# File 'lib/wp2txt/corpus.rb', line 69 def self.for_lang(lang, cache_dir: nil) manager = DumpManager.new(lang, cache_dir: cache_dir) multistream = manager.cached_multistream_path index = manager.cached_index_path unless File.exist?(multistream) && File.exist?(index) raise ArgumentError, "No cached dump for '#{lang}'. Run: wp2txt --build-index -L #{lang}" end new(multistream_path: multistream, index_path: index, cache_dir: cache_dir) end |
Instance Method Details
#category_tree(category, depth: 2) ⇒ Object
201 202 203 |
# File 'lib/wp2txt/corpus.rb', line 201 def category_tree(category, depth: 2) { dump: dump_name, tree: @metadata.category_tree(category, depth: depth) } end |
#close ⇒ Object
434 435 436 |
# File 'lib/wp2txt/corpus.rb', line 434 def close close_read_connections end |
#describe_schema ⇒ Object
CREATE statements of all tables available to query_sql
425 426 427 428 429 430 431 432 |
# File 'lib/wp2txt/corpus.rb', line 425 def describe_schema db = readonly_db schemas = { meta: db.execute("SELECT sql FROM sqlite_master WHERE sql IS NOT NULL").map(&:first) } if fts.built? schemas[:fts] = db.execute("SELECT sql FROM fts.sqlite_master WHERE sql IS NOT NULL").map(&:first) end schemas end |
#dump_info ⇒ Object
100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 |
# File 'lib/wp2txt/corpus.rb', line 100 def dump_info stats = @metadata.built? ? @metadata.stats : nil { multistream_path: @multistream_path, dump: stats&.dig(:dump_name) || File.basename(@multistream_path), tiers: { titles: File.exist?(@index_path), metadata: @metadata.built?, fulltext: fts.built? }, metadata_current: @metadata.built? && @metadata.valid_for?(@multistream_path), fulltext_current: fts.built? && fts.valid_for?(@multistream_path), stats: stats, fulltext: fts.built? ? fts.stats : nil } end |
#extract_corpus(output_path:, content: "sections", sections: nil, alias_set: nil, category: nil, depth: 0, categories: nil, category_match: nil, title_match: nil, limit: 0, chunk_size: nil, chunk_overlap: 0, max_articles: DEFAULT_MAX_SYNC_ARTICLES, num_processes: 4, progress: nil, cancel_check: nil) ⇒ Object
310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 |
# File 'lib/wp2txt/corpus.rb', line 310 def extract_corpus(output_path:, content: "sections", sections: nil, alias_set: nil, category: nil, depth: 0, categories: nil, category_match: nil, title_match: nil, limit: 0, chunk_size: nil, chunk_overlap: 0, max_articles: DEFAULT_MAX_SYNC_ARTICLES, num_processes: 4, progress: nil, cancel_check: nil) if content == "sections" && Array(sections).empty? && alias_set.nil? raise ArgumentError, "content: \"sections\" requires sections or alias_set" end raise ArgumentError, "chunk_overlap must be smaller than chunk_size" if chunk_size && chunk_overlap >= chunk_size raise ArgumentError, "chunking is not supported for content: \"wikitext\"" if chunk_size && content == "wikitext" filters = { category: category, depth: depth, categories: categories, category_match: category_match, sections: sections, alias_set: alias_set, title_match: title_match } total = @metadata.count_articles(**filters) cap = if limit.positive? max_articles ? [limit, max_articles].min : limit else max_articles || total end titles = @metadata.find_articles(**filters, limit: cap) truncated = total > titles.size resolved_sections = content == "summary" ? [SectionExtractor::SUMMARY_KEY] : (sections, alias_set) alias_contents = alias_set ? get_alias_set(alias_set)&.dig(:groups) : nil # Close ALL SQLite connections before MultistreamReader forks workers # (children must not inherit open database handles) close_read_connections articles_extracted = 0 records_written = 0 titles_done = 0 sample = [] File.open(output_path, "w") do |f| titles.each_slice(EXTRACT_BATCH_SIZE) do |batch| raise Cancelled if cancel_check&.call pages = reader.extract_articles_parallel(batch, num_processes: num_processes) batch.each do |t| page = pages[t] next unless page records = build_records(page, content, resolved_sections, chunk_size, chunk_overlap) next if records.empty? articles_extracted += 1 records.each do |record| f.puts(JSON.generate(record)) records_written += 1 sample << record if sample.size < 3 end end titles_done += batch.size progress&.call(titles_done, titles.size) end end = "#{output_path}.meta.json" File.write(, JSON.pretty_generate( tool: "wp2txt #{Wp2txt::VERSION}", dump: dump_name, generated_at: Time.now.utc.iso8601, query: filters.compact.merge(content: content, resolved_sections: resolved_sections, chunk_size: chunk_size, chunk_overlap: chunk_size ? chunk_overlap : nil).compact, alias_set_contents: alias_contents, total_matching: total, articles_extracted: articles_extracted, records_written: records_written, truncated: truncated )) { output_path: output_path, meta_path: , dump: dump_name, total_matching: total, articles_extracted: articles_extracted, records_written: records_written, truncated: truncated, bytes: File.size(output_path), sample: sample } end |
#find_articles(**filters) ⇒ Object
Tier 1: exhaustive queries (delegated to MetadataIndex)
193 194 195 196 197 198 199 |
# File 'lib/wp2txt/corpus.rb', line 193 def find_articles(**filters) limit = filters.delete(:limit) || 0 offset = filters.delete(:offset) || 0 total = @metadata.count_articles(**filters) titles = @metadata.find_articles(**filters, limit: limit, offset: offset) { dump: dump_name, total: total, returned: titles.size, titles: titles } end |
#fts ⇒ Object
117 118 119 120 121 122 |
# File 'lib/wp2txt/corpus.rb', line 117 def fts @fts ||= FtsIndex.new( FtsIndex.path_for(@multistream_path, cache_dir: @cache_dir), @metadata.db_path ) end |
#get_alias_set(name) ⇒ Object
284 285 286 |
# File 'lib/wp2txt/corpus.rb', line 284 def get_alias_set(name) @metadata.get_alias_set(name) end |
#get_article(title, format: "text", follow_redirect: true, max_chars: DEFAULT_MAX_CHARS) ⇒ Object
135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 |
# File 'lib/wp2txt/corpus.rb', line 135 def get_article(title, format: "text", follow_redirect: true, max_chars: DEFAULT_MAX_CHARS) page = fetch_page(title, follow_redirect: follow_redirect) return nil unless page body = case format.to_s when "wikitext" page[:text] else render_text(page) end result = { id: page[:id], title: page[:title], format: format.to_s } if max_chars && body.length > max_chars result.merge(text: body[0, max_chars], truncated: true, total_chars: body.length) else result.merge(text: body) end end |
#get_categories(title) ⇒ Object
Categories of one article, resolved through the same title normalization
154 155 156 157 158 159 160 161 162 163 |
# File 'lib/wp2txt/corpus.rb', line 154 def get_categories(title) cats = @metadata.categories_of(title) if cats.nil? page = fetch_page(title) cats = page ? @metadata.categories_of(page[:title]) : nil end return nil unless cats { title: title, categories: cats } end |
#get_sections(title, sections, alias_set: nil) ⇒ Object
Extract specific sections from one article.
168 169 170 171 172 173 174 175 176 177 178 |
# File 'lib/wp2txt/corpus.rb', line 168 def get_sections(title, sections, alias_set: nil) page = fetch_page(title) return nil unless page resolved = (sections, alias_set) config = RENDER_CONFIG.merge(format: :json, sections: resolved, title: page[:title]) article = Article.new(page[:text], page[:title], false) result = format_with_sections(article, config) { id: page[:id], title: page[:title], requested: sections, resolved: resolved, sections: result ? result["sections"] : {} } end |
#list_alias_sets ⇒ Object
288 289 290 |
# File 'lib/wp2txt/corpus.rb', line 288 def list_alias_sets @metadata.list_alias_sets end |
#list_headings(title) ⇒ Object
180 181 182 183 184 185 186 187 |
# File 'lib/wp2txt/corpus.rb', line 180 def list_headings(title) page = fetch_page(title) return nil unless page article = Article.new(page[:text], page[:title], false) { id: page[:id], title: page[:title], headings: SectionExtractor.new.extract_headings_with_levels(article) } end |
#metadata_built? ⇒ Boolean
Info
96 97 98 |
# File 'lib/wp2txt/corpus.rb', line 96 def @metadata.built? end |
#query_sql(sql, limit: SQL_ROW_LIMIT, timeout: SQL_TIMEOUT_SECONDS) ⇒ Object
Run a read-only SELECT against the metadata DB (with the FTS DB attached
as fts when built). Defense layers: keyword screening (outside string
literals), an SQLITE_OPEN_READONLY connection so writes are impossible at
the driver level, and subprocess execution with a hard wall-clock timeout
— the sqlite3 gem holds the GVL during C execution, so a runaway query
can only be stopped by killing the process running it. On timeout, the
error message includes an EXPLAIN QUERY PLAN diagnosis when a likely
cause (nested full scans, unbounded recursion) is recognizable.
407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 |
# File 'lib/wp2txt/corpus.rb', line 407 def query_sql(sql, limit: SQL_ROW_LIMIT, timeout: SQL_TIMEOUT_SECONDS) raise ArgumentError, "only SELECT/WITH queries are allowed" unless sql =~ /\A\s*(SELECT|WITH)\b/i # Screen keywords outside string literals and quoted identifiers only, so # legitimate data values (e.g. title LIKE '%Update%') are not rejected screened = sql.gsub(/'(?:[^']|'')*'/m, "''").gsub(/"(?:[^"]|"")*"/m, '""') raise ArgumentError, "query contains a forbidden keyword" if screened.match?(SQL_FORBIDDEN) limit = [[limit.to_i, 1].max, 1000].min if Process.respond_to?(:fork) run_sql_in_subprocess(sql, limit, timeout) else run_sql_on(readonly_db, sql, limit) end rescue SQLite3::Exception => e raise ArgumentError, "SQL error: #{e.}" end |
#save_alias_set(name, groups, force: false, max_ratio: GUARDRAIL_MAX_RATIO, min_articles: GUARDRAIL_MIN_ARTICLES) ⇒ Object
261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 |
# File 'lib/wp2txt/corpus.rb', line 261 def save_alias_set(name, groups, force: false, max_ratio: GUARDRAIL_MAX_RATIO, min_articles: GUARDRAIL_MIN_ARTICLES) unless groups.is_a?(Array) && !groups.empty? && groups.all? { |g| g.is_a?(Array) && !g.empty? } raise ArgumentError, "groups must be a non-empty array of arrays" end # Report the exact thresholds used so both the model and the user can see # why a group was accepted or rejected, not just that it was criteria = { max_cooccurrence_ratio: max_ratio, min_articles: min_articles } violations = check_alias_groups(groups, max_ratio: max_ratio, min_articles: min_articles) if violations.any? && !force return { saved: false, name: name, groups: groups, violations: violations, criteria: criteria, warning: "These heading pairs coexist in the same article more than " \ "#{(max_ratio * 100).round}% of the time (both headings appear in at least " \ "#{min_articles} articles), so they are likely different sections, not synonyms. " \ "Remove them from the group, or pass force: true if you verified them another " \ "way (e.g., by reading section contents with get_sections)." } end @metadata.save_alias_set(name, groups) { saved: true, name: name, groups: groups, violations: violations, criteria: criteria } end |
#search_text(query, mode: "phrase", sections: nil, alias_set: nil, category: nil, depth: 0, limit: 20, offset: 0, count: "capped", snippets: true) ⇒ Object
Exhaustive full-text search over cleaned section text.
235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 |
# File 'lib/wp2txt/corpus.rb', line 235 def search_text(query, mode: "phrase", sections: nil, alias_set: nil, category: nil, depth: 0, limit: 20, offset: 0, count: "capped", snippets: true) unless fts.built? raise ArgumentError, "Full-text index not built. Run: wp2txt --build-index --fulltext" end resolved = (sections, alias_set) resolved = nil if resolved.empty? result = fts.search(query, mode: mode, sections: resolved, category: category, depth: depth, limit: limit, offset: offset, count: count) hits = result[:hits] attach_snippets(hits, query, mode) if snippets && !hits.empty? { dump: dump_name, query: query, mode: mode, total: result[:total], total_is_capped: result[:total_is_capped], returned: hits.size, hits: hits.map do |h| { page_id: h[:page_id], title: h[:title], section: h[:heading].to_s.empty? ? nil : h[:heading], section_path: h[:heading].to_s.empty? ? h[:title] : "#{h[:title]} > #{h[:heading]}", snippet: h[:snippet] }.compact end } end |
#section_cooccurrence(headings, category: nil, depth: 0) ⇒ Object
211 212 213 214 |
# File 'lib/wp2txt/corpus.rb', line 211 def section_cooccurrence(headings, category: nil, depth: 0) @metadata.section_cooccurrence(headings, category: category, depth: depth) .merge(dump: dump_name) end |
#section_stats(category: nil, depth: 0, top_n: 50) ⇒ Object
205 206 207 208 209 |
# File 'lib/wp2txt/corpus.rb', line 205 def section_stats(category: nil, depth: 0, top_n: 50) { dump: dump_name, sections: @metadata.section_stats(category: category, depth: depth, top_n: top_n) .map { |h, c| { heading: h, articles: c } } } end |