PdfDocument.pdf_spans
def pdf_spans(
doc:PdfDocument, st:int=0, end:NoneType=None
)->L:Extract text spans with font metadata (size, weight, bbox) per page.
We will build a simple ingestion pipeline to ingest pdf documents into litesearch database for searching.
Extensions to pdf_oxide PdfDocument to extract texts, images, links, tables, spans and outline
Extract text spans with font metadata (size, weight, bbox) per page.
Extract structured tables (rows/cells/bbox) from each page.
Convert pages to markdown with headings and tables detected.
Extract images — returns metadata dicts, or saves to output_dir if provided.
Extract hyperlinks from each page as a list of URI strings.
Extract plain text from each page.
The PDF helpers are pdflite, re-exported here.
pdf_parse returns one markdown string per page. The document tree and the measured chunk sizes are built on that split.
def pdf_parse(
pdf:pdf_oxide.pdf_oxide.PdfDocument | str | pathlib.Path | bytes,
out_path:str | pathlib.Path=None, # dir holding the image subdir; links are relative to it
image_dir:str='images', # subdir for extracted images
preserve_layout:bool=False, ocr_selection:str='auto', # auto | on | off
)->list:A PDF as one markdown string per page. The page split the tree and its chunk sizes are measured against.
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
SafeFastChunkerFastChunker whose cuts land on character boundaries.
chonkie-core returns byte offsets and FastChunker.chunk decodes each slice on its own, so a size cut that falls inside a multi-byte character raises UnicodeDecodeError, which any long unbroken paragraph containing an em dash or a curly quote will do, i.e. most scraped web pages. Both sides of a cut are moved back to the start of the character it split, so the text is repartitioned by up to three bytes and never lost.
FastChunker whose cuts land on character boundaries.
Markdown-chunked text from each page; yields (page, chunk_idx, text) tuples.
Split markdown into paragraph chunks on blank lines
_pdf2 = next((p for p in ['pdfs/attention_is_all_you_need.pdf', 'nbs/pdfs/attention_is_all_you_need.pdf'] if Path(p).exists()), None)
if _pdf2:
_doc = PdfDocument(_pdf2)
_chunks = _doc.pdf_chunks()
assert len(_chunks) > 0
assert all(len(t) == 3 for t in _chunks), 'each chunk must be (page, chunk_idx, text)'
assert all(len(t[2]) >= 40 for t in _chunks), 'short fragments should be filtered'
print(f'pdf_chunks: {len(_chunks)} chunks from {_doc.page_count()} pages (sample: {repr(_chunks[0][2][:60])})...')
else:
print('Skipping chunk test — PDF not found')Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
pdf_chunks: 98 chunks from 15 pages (sample: 'Provided proper attribution is provided, Google hereby grant')...
Split text into chunks, returning (start_char, end_char, text) spans into the original text.
for _bad in ['a'*510 + '—' + 'b'*600, # the cut lands mid-character
'a'*400 + '—'*200, # every candidate cut is mid-character
'—'*700]:
assert ''.join(chunk_markdown(_bad)) == _bad, 'text lost or altered by chunking'
for s, e, txt in chunk_spans(_bad): assert _bad[s:e] == txt
# the upstream chunker on the same input is what used to reach `add_doc`
try: FastChunker(chunk_size=512)('a'*510 + '—' + 'b'*600); assert False, 'chonkie fixed it upstream'
except UnicodeDecodeError: pass
# a paragraph split still happens where it should, and ASCII text chunks exactly as before
_para = ('x'*200 + '\n\n')*3
assert chunk_markdown(_para, SafeFastChunker(chunk_size=512, delimiters='\n\n')) == \
chunk_markdown(_para, FastChunker(chunk_size=512, delimiters='\n\n'))Code extraction utilities
Signatures for paths that are not python code files using codesigs
Parse a code string or python file and return code chunks as list of dicts with content and metadata.
Only module-level definitions are ever returned, so tree.body is the candidate list: walking the whole tree to tag every node with a parent and then filtering on “parent is the Module” computes an answer the body already gives. Dropping the walk and hoisting the line split take the same 400 files from 23.4s to 3.3s, chunk-for-chunk identical.
Return the importlib ModuleSpec for a package, or None if not found.
Find the root of the current git repository, or None if not in a repo.
add_doc(chunker=...) implies that a format is described by its chunker. It is not. Reading Sanskrit correctly needed three different things: a parser that knows <lg xml:id=…> is a verse, a tree mode that knows a citation names a chapter, and a chunker that cuts on daṇḍas. A Profile names all three together, so the next format is a registration rather than another branch inside add_file.
Extension is a filter, never the decision, .xml is TEI, vedicreader and a hundred unrelated schemas, so a profile sharing an extension must sniff content.
ProfileHow one kind of document should be read: parser, structure, chunker.
A chunker alone is not enough to describe a format, which is what add_doc(chunker=...) on its own implies. Getting Sanskrit right needed a reader that knows <lg xml:id=…> is a verse, a tree mode that knows a citation is a chapter boundary, and a chunker that cuts on daṇḍas, three different seams. A profile is the one object that names all three, so adding the next format is a registration rather than another branch in add_file.
meta is the fourth: what a format knows about a chunk that no generic pipeline could compute. Sanskrit fills it with metre and gaṇa signature, which land in the metadata column get_store already indexes for FTS, so a facet nothing else in litesearch understands becomes searchable and filterable with no schema change.
profile_forThe profile that claims this document, or None for the default pipeline.
Extension is a filter, never the decision: .xml is TEI, vedicreader and a hundred unrelated schemas, and .txt is anything at all. A profile that shares an extension must sniff the content, and one that cannot is only reached by name.
Parse a code file or code string and return code chunks with metadata.
The profile that claims this document, or None for the default pipeline.
Add a profile to the registry, replacing any of the same name.
How one kind* of document should be read: parser, structure, chunker.*
{'content': '---\ndescription: Building blocks for litesearch\noutput-file: core.html\ntitle: core\n\n---\n\n',
'metadata': {'path': '01_core.ipynb',
'uploaded_at': 1787911231.2660935,
'lang': '.ipynb',
'type': 'raw'}}
You can use pyparse to extract code chunks from a python file or code string.
[{'content': 'class SomeClass:\n def __init__(self,x): store_attr()\n def method(self): return self.x + a', 'metadata': {'path': 'None', 'uploaded_at': None, 'name': 'SomeClass', 'lang': '.py', 'type': 'ClassDef', 'lineno': 4, 'end_lineno': 6}}]
Setting imports to True will also include import statements as code chunks.
[{'content': 'from fastcore.all import *', 'metadata': {'path': 'None', 'uploaded_at': None, 'name': None, 'lang': '.py', 'type': 'ImportFrom', 'lineno': 2, 'end_lineno': 2}}, {'content': 'class SomeClass:\n def __init__(self,x): store_attr()\n def method(self): return self.x + a', 'metadata': {'path': 'None', 'uploaded_at': None, 'name': 'SomeClass', 'lang': '.py', 'type': 'ClassDef', 'lineno': 4, 'end_lineno': 6}}]
[{'content': 'a=1', 'metadata': {'path': 'None', 'uploaded_at': None, 'name': None, 'lang': '.py', 'type': 'Assign', 'lineno': 3, 'end_lineno': 3}}, {'content': 'class SomeClass:\n def __init__(self,x): store_attr()\n def method(self): return self.x + a', 'metadata': {'path': 'None', 'uploaded_at': None, 'name': 'SomeClass', 'lang': '.py', 'type': 'ClassDef', 'lineno': 4, 'end_lineno': 6}}]
def parse_js():
from textwrap import dedent
from tempfile import NamedTemporaryFile as ntf
js_code = dedent("""
const htmlElement = document.documentElement;
const mediaQuery = window.matchMedia('(prefers-color-scheme: dark)');
const vr = '__VR__';
let __VR__ = JSON.parse(localStorage.getItem(vr) || '{{__state__}}');
function storeState(key, value) {__VR__[key] = value; localStorage.setItem(vr, JSON.stringify(__VR__));}
function getState(key){return __VR__[key];}
function removeState(key) {delete __VR__[key]; localStorage.setItem(key, JSON.stringify(__VR__)); }
function setTheme(color,fn=null, ...args) {
if (color === null || color === undefined) {return;}
htmlElement.classList.remove(__VR__.theme);
htmlElement.classList.add(color);
storeState('theme', color);
if (typeof fn === 'function') {fn(...args);}
}
""")
with ntf('w+', suffix='.js', encoding='utf-8') as f:
f.write(js_code)
f.flush() # Ensure content is readable by file_parse from disk
print(file_parse(Path(f.name)))
parse_js()[{'content': 'function storeState(key, value) {...}', 'metadata': {'path': '/tmp/tmp9wvads4o.js', 'uploaded_at': 1788044982.810499, 'lang': '.js'}}, {'content': 'function getState(key) {...}', 'metadata': {'path': '/tmp/tmp9wvads4o.js', 'uploaded_at': 1788044982.810499, 'lang': '.js'}}, {'content': 'function removeState(key) {...}', 'metadata': {'path': '/tmp/tmp9wvads4o.js', 'uploaded_at': 1788044982.810499, 'lang': '.js'}}, {'content': 'function setTheme(color,fn=null, ...args) {...}', 'metadata': {'path': '/tmp/tmp9wvads4o.js', 'uploaded_at': 1788044982.810499, 'lang': '.js'}}]
def pkg2files(
pkg:str, # package name
skip_file_re:str='^[.]|^(?:setup\\.py|conftest\\.py)$', # regex to skip files
skip_folder_re:str='^[.]|^(?:tests?|examples?|docs?|build|dist|_proc|_docs|_site|_site_libs|node_modules|__pycache__)$', # regex to skip folders
func:type=Path, # function to apply to file paths (e.g. Path or str)
path:pathlib.Path | str='.', # path to start searching
recursive:bool=True, # search subfolders
maxdepth:int=None, # max depth to descend (1=just immediate contents; None=unlimited)
symlinks:bool=True, # follow symlinks?
file_glob:str=None, # Only include files matching glob
file_re:str=None, # Only include files matching regex
folder_re:str=None, # Only enter folders matching regex
skip_file_glob:str=None, # Skip files matching glob
ret_folders:bool=False, # return folders, not just files
sort:bool=True, # sort files by name within each folder
types:str | list=None, # list or comma-separated str of ext types from: py, js, java, c, cpp, rb, r, ex, sh, web, doc, cfg
exts:str | list=None, # list or comma-separated str of exts to include
)->L: # additional args to pass to globtasticReturn list of python files in a package excluding tests and setup files.
[Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/fastlite/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/fastlite/_modidx.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/fastlite/core.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/fastlite/kw.py')]
[Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/_upb/_message.abi3.so'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/any.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/any_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/api_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/compiler/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/compiler/plugin_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/descriptor.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/descriptor_database.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/descriptor_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/descriptor_pool.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/duration.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/duration_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/empty_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/field_mask_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/api_implementation.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/builder.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/containers.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/decoder.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/encoder.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/enum_type_wrapper.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/extension_dict.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/field_mask.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/message_benchmark.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/message_listener.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/python_edition_defaults.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/python_message.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/testing_refleaks.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/type_checkers.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/well_known_types.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/internal/wire_format.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/json_enumvalue_options_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/json_format.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/json_options_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/message.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/message_factory.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/proto.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/proto_builder.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/proto_json.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/proto_text.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/pyext/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/pyext/cpp_message.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/reflection.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/runtime_version.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/service_reflection.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/source_context_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/struct_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/symbol_database.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/testdata/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/text_encoding.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/text_format.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/timestamp.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/timestamp_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/type_pb2.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/unknown_fields.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/util/__init__.py'), Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/google/protobuf/wrappers_pb2.py')]
def dir2files(
dir:str, # package name
skip_file_re:str='^[.]|^(?:setup\\.py|conftest\\.py)$', # regex to skip files
skip_folder_re:str='^[.]|^(?:tests?|examples?|docs?|build|dist|_proc|_docs|_site|_site_libs|node_modules|__pycache__)$', # regex to skip folders
func:type=Path, # function to apply to file paths (e.g. Path or str)
path:pathlib.Path | str='.', # path to start searching
recursive:bool=True, # search subfolders
maxdepth:int=None, # max depth to descend (1=just immediate contents; None=unlimited)
symlinks:bool=True, # follow symlinks?
file_glob:str=None, # Only include files matching glob
file_re:str=None, # Only include files matching regex
folder_re:str=None, # Only enter folders matching regex
skip_file_glob:str=None, # Skip files matching glob
ret_folders:bool=False, # return folders, not just files
sort:bool=True, # sort files by name within each folder
types:str | list=None, # list or comma-separated str of ext types from: py, js, java, c, cpp, rb, r, ex, sh, web, doc, cfg
exts:str | list=None, # list or comma-separated str of exts to include
)->L: # additional args to pass to globtasticReturn list of python files in a package excluding tests and setup files.
[Path('/Users/71293/code/litesearch/evals/__init__.py'), Path('/Users/71293/code/litesearch/evals/boldhead.py'), Path('/Users/71293/code/litesearch/evals/build.py'), Path('/Users/71293/code/litesearch/evals/cluster_eval.py'), Path('/Users/71293/code/litesearch/evals/corpus.py'), Path('/Users/71293/code/litesearch/evals/decide.py'), Path('/Users/71293/code/litesearch/evals/encoders.py'), Path('/Users/71293/code/litesearch/evals/extractor_eval.py'), Path('/Users/71293/code/litesearch/evals/extractor_sig.py'), Path('/Users/71293/code/litesearch/evals/fastlate.py'), Path('/Users/71293/code/litesearch/evals/fetch_corpus.py'), Path('/Users/71293/code/litesearch/evals/ingest_bench.py'), Path('/Users/71293/code/litesearch/evals/latechunk.py'), Path('/Users/71293/code/litesearch/evals/multihop.py'), Path('/Users/71293/code/litesearch/evals/queries.py'), Path('/Users/71293/code/litesearch/evals/refindex.py'), Path('/Users/71293/code/litesearch/evals/report.py'), Path('/Users/71293/code/litesearch/evals/run.py'), Path('/Users/71293/code/litesearch/evals/score.py'), Path('/Users/71293/code/litesearch/evals/tables.py'), Path('/Users/71293/code/litesearch/litesearch/__init__.py'), Path('/Users/71293/code/litesearch/litesearch/_modidx.py'), Path('/Users/71293/code/litesearch/litesearch/api.py'), Path('/Users/71293/code/litesearch/litesearch/cli.py'), Path('/Users/71293/code/litesearch/litesearch/core.py'), Path('/Users/71293/code/litesearch/litesearch/data.py'), Path('/Users/71293/code/litesearch/litesearch/postfix.py'), Path('/Users/71293/code/litesearch/litesearch/quality.py'), Path('/Users/71293/code/litesearch/litesearch/sanskrit.py'), Path('/Users/71293/code/litesearch/litesearch/topics.py'), Path('/Users/71293/code/litesearch/litesearch/tree.py'), Path('/Users/71293/code/litesearch/litesearch/utils.py')]
def dir2chunks(
dir:str, # directory path
imports:bool=False, # include import statements as code chunks
assigns:bool=False, # include top-level assignments as code chunks
n_workers:int=None, # parse workers; 0 is serial, None picks by file count
threadpool:bool=False, # threads rather than processes (see `_parse_files`)
skip_file_re:str='^[.]|^(?:setup\\.py|conftest\\.py)$', # regex to skip files
skip_folder_re:str='^[.]|^(?:tests?|examples?|docs?|build|dist|_proc|_docs|_site|_site_libs|node_modules|__pycache__)$', # regex to skip folders
func:type=Path, # function to apply to file paths (e.g. Path or str)
path:pathlib.Path | str='.', # path to start searching
recursive:bool=True, # search subfolders
maxdepth:int=None, # max depth to descend (1=just immediate contents; None=unlimited)
symlinks:bool=True, # follow symlinks?
file_glob:str=None, # Only include files matching glob
file_re:str=None, # Only include files matching regex
folder_re:str=None, # Only enter folders matching regex
skip_file_glob:str=None, # Skip files matching glob
ret_folders:bool=False, # return folders, not just files
sort:bool=True, # sort files by name within each folder
types:str | list=None, # list or comma-separated str of ext types from: py, js, java, c, cpp, rb, r, ex, sh, web, doc, cfg
exts:str | list=None, # list or comma-separated str of exts to include
)->L: # additional args to pass to dir2filesReturn code chunks from a directory with extra metadata.
def pkg2chunks(
pkg:str, # package name
imports:bool=False, # include import statements as code chunks
n_workers:int=None, # parse workers; 0 is serial, None picks by file count
threadpool:bool=False, # threads rather than processes (see `_parse_files`)
**kw
)->L: # additional args to pass to pkg2filesReturn code chunks from a package with extra metadata.
pkg2chunks can be used to extract code chunks from an entire package installed in your environment.
{'content': 'def t(self:Database): return _TablesGetter(self)',
'metadata': {'path': '/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/fastlite/core.py',
'uploaded_at': 1773452878.5692947,
'name': 't',
'lang': '.py',
'type': 'FunctionDef',
'lineno': 44,
'end_lineno': 44,
'package': 'fastlite',
'version': '0.2.4'}}
Path('/Users/71293/code/litesearch/.venv/lib/python3.13/site-packages/pdf_oxide/__init__.py')
{'content': '---\ndescription: Fixes for post import issues\noutput-file: postfix.html\ntitle: postfix\n\n---\n\n',
'metadata': {'path': '../nbs/00_postfix.ipynb',
'uploaded_at': 1773099020.1621501,
'lang': '.ipynb',
'type': 'raw',
'dir': '../nbs'}}
Return list of installed packages. If nms is provided, return only those packages.
Get list of installed packages in your environment using installed_packages. If you pass a list of package names, it only returns them if they exist in your environment.
['pdflite', 'chonkie', 'pillow', 'model2vec', 'tokenizers', 'pdf_oxide', 'fastlite', 'usearch', 'codesigs']
_ls = spec('litesearch')
assert _ls and spec('litesearch[sanskrit]')
assert spec('litesearch[sanskrit]').name == _ls.name
assert 'litesearch' in installed_packages(['litesearch[sanskrit]'])
_rishi = spec('rishi')
if _rishi:
assert spec('rishi[all]')
assert spec('rishi[all]').name == _rishi.name
assert 'rishi' in installed_packages(['rishi[all]'])['shellingham', 'nbdev', 'python-fastllm', 'tornado', 'selectolax', 'threadpoolctl', 'apswutils', 'uri-template', 'rfc3339-validator', 'pexpect']
['chonkie', 'codesigs', 'fastlite', 'model2vec', 'tokenizers', 'pdf_oxide', 'pdflite', 'pillow', 'usearch']
['chonkie', 'codesigs', 'fastlite', 'model2vec', 'tokenizers', 'pdf_oxide', 'pdflite', 'pillow', 'usearch', 'nbdev', 'jupyterlab', 'jupyterlab-quarto', 'notebook', 'FlashRank', 'onnx', 'onnxruntime', 'model2vec', 'datasets', 'rishi', 'koshas', 'fastprogress', 'selectolax', 'pyperclip']
Query Preprocessing utilities
cleanClean the query by removing * and returning None for empty queries.
_ is deliberately kept.
add_wcAdd wild card * to each word in the query.kw uses UAX#29 segmentation, which keeps
the apostrophe inside pathfinder's,
Preprocess the query for fts search.
Query keywords by UAX#29 word segmentation (apsw), minus stopwords.
Widen the query by joining words with OR operator.
Add wild card to each word in the query.kw uses UAX#29 segmentation, which keeps*
Clean the query by removing and returning None for empty queries.*
You can clean queries passed into fts search using clean, add wild cards using add_wc, widen the query using mk_wider and extract keywords using kw. You can combine all these using pre function.
q = 'This is a sample query'
print('preprocessed q with defaults: `%s`' %pre(q))
print('keywords extracted: `%s`' %pre(q, wc=False, wide=False))
print('q with wild card: `%s`' %pre(q, extract_kw=False, wide=False, wc=True))
print('remove _ : `%s`' %pre('This_is_a_query_with_underscores with_uv', extract_kw=False))preprocessed q with defaults: `sample* OR query*`
keywords extracted: `sample query`
q with wild card: `This* is* a* sample* query*`
remove _ : `This_is_a_query_with_underscores* OR with_uv*`
# A possessive used to make `pre` emit invalid FTS5, so `search` raised instead of searching.
# `pathfinder's*` is a syntax error; `"pathfinder's"*` is a prefix query on a string.
assert pre("The Pathfinder's transition") == '"pathfinder\'s"* OR transition*'
import apsw
_c = apsw.Connection(':memory:')
_c.execute("create virtual table t using fts5(content)")
_c.execute("insert into t values(?)", ("The Pathfinder's transition activities",))
_c.execute("insert into t values(?)", ('something else entirely',))
_q = pre("Pathfinder's activities")
assert len(list(_c.execute("select content from t where t match ?", (_q,)))) == 1img2png normalises any image input (PIL, bytes, or path) to RGB PNG bytes. png_det parses the PNG header and IDAT stream. images_to_pdf wraps a list of images into a conformant multi-page image-only PDF without any external dependencies.
Note:
pdf_oxide.Pdf.from_image*methods are broken in 0.3.x (produce an empty 1.3 KB shell). These functions build a conformant PDF directly from the PNG spec using FlateDecode + Predictor 15.
Convert a list of images into a multi-page image-only PDF (one image per page).
Return (width, height, concatenated IDAT bytes) from a PNG.
Normalise image input (PIL Image / bytes / path) to PNG bytes as RGB.
# find PDF relative to nbs/ (normal Jupyter) or project root (nbdev_test)
_pdf = next((p for p in ['pdfs/attention_is_all_you_need.pdf', 'nbs/pdfs/attention_is_all_you_need.pdf'] if Path(p).exists()), None)
assert _pdf, 'attention_is_all_you_need.pdf not found'
arxiv = PdfDocument(_pdf)
# pdf_texts and pdf_markdown
md_pages = arxiv.pdf_markdown()
assert len(md_pages) == arxiv.page_count(), f'Expected {arxiv.page_count()} pages, got {len(md_pages)}'
assert any('attention' in p.lower() for p in md_pages)
txt_pages = arxiv.pdf_texts()
assert len(txt_pages) == arxiv.page_count()
assert any('attention' in t.lower() for t in txt_pages)
links = arxiv.pdf_links()
assert isinstance(links, L)
print(f'pdf_texts: {len(txt_pages)} pages, pdf_markdown: {len(md_pages)} pages, pdf_links: {len(links)} links')
# images_to_pdf: build synthetic scanned PDF from a real page image
with tempfile.TemporaryDirectory() as tmp:
arxiv.to_markdown(2, include_images=True, embed_images=False, image_output_dir=tmp)
img_bytes = open(os.path.join(tmp, os.listdir(tmp)[0]), 'rb').read()
_scanned = Path(_pdf).parent/'scanned_test.pdf'
images_to_pdf([img_bytes], output=str(_scanned))
scanned = PdfDocument(str(_scanned))
assert scanned.extract_text(0).strip() == '', 'scanned PDF should have no native text'
assert _scanned.stat().st_size > 100_000, 'scanned PDF should contain the embedded image'
print('images_to_pdf test passed:', _scanned)Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
Dictionary used where Stream expected, treating as empty stream
pdf_texts: 15 pages, pdf_markdown: 15 pages, pdf_links: 18 links
images_to_pdf test passed: pdfs/scanned_test.pdf
Copy bundled SKILL.md to skill directories (.agents/, .claude/, .Codex/).
import tempfile, importlib.util as _iu
from pathlib import Path as _Path
# Locate the installed litesearch package — works regardless of cwd
_spec = _iu.find_spec('litesearch')
skill_src = _Path(_spec.submodule_search_locations[0]) / 'SKILL.md'
assert skill_src.exists(), f'SKILL.md not found at {skill_src.resolve()}'
with tempfile.TemporaryDirectory() as tmp:
tmp = _Path(tmp)
for d in ['.agents', '.claude', '.Codex']:
dest = tmp / d / 'skills' / 'litesearch' / 'SKILL.md'
dest.parent.mkdir(parents=True, exist_ok=True)
dest.write_text(skill_src.read_text())
for d in ['.agents', '.claude', '.Codex']:
dest = tmp / d / 'skills' / 'litesearch' / 'SKILL.md'
assert dest.exists(), f'Missing: {dest}'
assert 'litesearch' in dest.read_text()
print('mv_skill_md install paths ok')mv_skill_md install paths ok