Analyze Text Helper¶
openmed.analyze_text is the top-level orchestrator that most users start with. It validates input, spins up a token-classification pipeline, segments sentences, and normalizes the output so you can copy dict/JSON/HTML/CSV payloads straight into downstream systems.
Quick reference¶
from openmed import analyze_text
result = analyze_text(
text="Patient started on imatinib for chronic myeloid leukemia.",
model_name="disease_detection_superclinical",
aggregation_strategy="simple",
output_format="dict",
include_confidence=True,
confidence_threshold=0.55,
group_entities=True,
metadata={"source": "clinic-note-42"},
)
print(result.model)
print(result.entities[:3])
payload = result.to_dict()
Key arguments¶
model_name: registry alias, full Hugging Face id, or local model directory. Useopenmed.list_models()if you need auto-discovery.model_id: alias formodel_name, supported for API-style callers that use model-id terminology.aggregation_strategy: forwarded to the HF pipeline.simple(default) yields grouped tokens;Nonekeeps raw tokens.output_format:"dict"(default, returnsAnalyzeResult),"json","html", or"csv".include_confidence&confidence_threshold: control the final payload; defaults keep all scores.group_entities: merge adjacent spans of the same label after formatting.formatter_kwargs: forwarded toopenmed.processing.format_predictions.assert_context: opt in to deterministic clinical assertion labels. Each entity receivesnegation,uncertainty,experiencer, andtemporalityunderentity.metadata["clinical_context"].- Sentence options (
sentence_detection,sentence_language,sentence_clean,sentence_segmenter,sentence_backend) wrap the sentence engine so each prediction carries the sentence span; disable them if latency matters more than helper metadata.
The context stage is disabled by default. Enable it for clinical entities that will flow into grounding, FHIR export, or problem-list review:
result = analyze_text(
"No evidence of pneumonia.",
assert_context=True,
)
print(result.entities[0].metadata["clinical_context"])
Optional YASBD sentence backend¶
The default sentence_backend="auto" path is unchanged: OpenMed uses its built-in Indic and Chinese segmenters where appropriate and pySBD elsewhere. YASBD is neither installed nor imported by a core OpenMed installation.
Install the experimental backend explicitly when you want to benchmark it on your own workload:
Then select it for either the low-level sentence API or analyze_text:
from openmed import analyze_text
from openmed.processing.sentences import segment_text
spans = segment_text(
"Patient is stable. Follow up tomorrow.",
language="en",
backend="yasbd",
)
result = analyze_text(
"Patient is stable. Follow up tomorrow.",
sentence_backend="yasbd",
)
The adapter preserves OpenMed's exact, contiguous source offsets and assigns inter-sentence whitespace to the preceding span, matching the existing span contract. YASBD remains an explicit opt-in because sentence boundaries can differ between engines; validate representative clinical and multilingual inputs before adopting it in production. If the extra is missing, explicitly selecting "yasbd" raises an installation error instead of silently changing behavior.
Chunking & truncation¶
result = analyze_text(
text=long_report,
model_name="pharma_detection_superclinical",
aggregation_strategy=None, # work with raw tokens
max_length=512, # forwarded to HF pipeline
truncation=True, # enforce length (default)
sentence_detection=False, # skip sentence detection to save ~2ms per note
sentence_backend="auto", # "auto" (default) or "yasbd" (experimental, faster)
)
When you need full-control over tokenizer behaviour:
- Pass
max_length/truncationviapipeline_kwargs. If you skip truncation, the helper sets the tokenizer max length to unlimited (0) so HF pipelines accept longer inputs. - Provide
batch_sizeornum_workersinpipeline_kwargsand they will be forwarded to the pipeline call but not to the constructor. - Enable medical token remapping with
OpenMedConfig(use_medical_tokenizer=True)to group outputs onto clinical-friendly tokens without changing the model tokenizer.
Loading from a local path¶
Pass an existing model directory to model_name or model_id when the model files are already present on disk:
import os
from openmed import OpenMedConfig, analyze_text
local_path = os.path.abspath("./models/OpenMed-NER-DiseaseDetect-SuperClinical-434M")
config = OpenMedConfig(device="cpu")
result = analyze_text(
"Patient presents with chronic myeloid leukemia and Type 2 diabetes.",
model_id=local_path,
config=config,
)
for entity in result.entities:
print(entity.text, entity.label)
legacy_payload = result.to_dict()
print(legacy_payload["model_name"])
When the identifier points to an existing local path, OpenMed asks Transformers to load with local_files_only=True by default. That keeps air-gapped deployments from validating or downloading the model from the Hugging Face Hub. If any required tokenizer, config, or weight file is missing, loading fails locally with the underlying Transformers error.
Streaming multiple texts¶
analyze_text is optimized for single inputs. For batch jobs, keep a ModelLoader instance around and reuse its pipelines:
from openmed import ModelLoader, format_predictions
loader = ModelLoader()
pipeline = loader.create_pipeline("disease_detection_superclinical")
for note in notes:
raw = pipeline(note, batch_size=16)
formatted = format_predictions(raw, note, model_name="Disease Detection")
print(formatted.entities[:3])
See ModelLoader & Pipelines for details on caching, GPU selection, and tokenizer reuse.
HTML/CSV rendering¶
html = analyze_text(
text,
model_name="oncology_detection_superclinical",
output_format="html",
formatter_kwargs={
"html_class": "openmed-highlights",
"tag_colors": {"CANCER": "#d97706"},
},
)
csv_rows = analyze_text(
text,
model_name="pharma_detection_superclinical",
output_format="csv",
)
The HTML formatter emits a ready-to-embed snippet for dashboards; CSV mode writes row strings (header + body). Both respect confidence_threshold and group_entities.
Validation behaviours¶
Behind the scenes analyze_text calls:
validate_input— trims whitespace and enforces max lengths.validate_model_name— normalizes registry aliases.- Sentence detection (
openmed.processing.sentences) — optional segmentation with language hints and a selectable backend. OutputFormatter— see Advanced NER & Output Formatting for available kwargs.
If you need custom validation or logging, inject your own OpenMedConfig or reuse a configured ModelLoader.