Conversation
Implement multiple strategies for pulling fulltext from publications: - BioC XML API (NCBI BioNLP) for cleanest structured text - Europe PMC for wider open access coverage - Unpaywall API for finding OA versions via DOI - Identifier conversion utilities (DOI <-> PMID <-> PMCID) Integrate strategies into PMIDSource and DOISource with fallback: 1. BioC XML (fastest, cleanest) 2. Europe PMC (broader coverage) 3. Traditional PMC XML/HTML Based on patterns from aurelian and artl-mcp repositories.
Add BioCFullTextProvider and EuropePMCFullTextProvider as proper FullTextProvider implementations following the existing architecture. - bioc.py: Fetches structured text from NCBI BioNLP BioC XML API - epmc.py: Fetches fulltext XML for OA articles from Europe PMC (complements epmc_preprint.py which handles preprints only) Remove fulltext_strategies.py which bypassed the provider chain by fetching fulltext inside source classes. The provider architecture uses _enrich_with_full_text() in reference_fetcher.py to call providers after sources return content. Revert doi.py and pmid.py to main branch versions - fulltext enrichment happens via the provider chain, not inside sources. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PYbfu1JPi9rr7v5XU6qbJj
|
Claude finished @cmungall's task in 1m 21s —— View job Review: Add BioC and Europe PMC full-text providers
Both providers follow the existing pattern closely: they use registry registration, the shared 🔴 Probable bug: the Europe PMC fullTextXML URL may be wrong
EUROPEPMC_FULLTEXT_URL = ".../rest/{source}/{id}/fullTextXML"
pmcid_clean = pmcid.replace("PMC", "")
url = EUROPEPMC_FULLTEXT_URL.format(source="PMC", id=pmcid_clean)
# -> .../rest/PMC/123456/fullTextXMLAs far as I know, the documented Europe PMC endpoint is Suggested fix: 🟠 Correctness risk: BioC output includes the reference list (and other non-body passages)
A smaller issue of the same kind is in 🟡 Integration and docs: the providers can be reached but nobody will find them
🟡 Code quality and project conventions
🟢 PerformanceThat's fine. 🔒 SecurityThere are no concerns. Inputs are PMIDs and PMCIDs placed in fixed HTTPS hosts. 🧪 Test coverageThe happy and sad paths are present, but these are missing:
Summary: the structure is solid. Before merging, I'd (1) check and fix the Europe PMC fullTextXML URL, (2) exclude |
Summary
Adds two new full-text providers to expand article text retrieval capabilities: BioC XML provider for structured paragraph-level text from PMC Open Access subset, and Europe PMC provider for fetching full-text XML from open access articles.
Key Changes
New BioC Provider (
src/linkml_reference_validator/etl/fulltext/bioc.py):FullTextProviderRegistryas "bioc"New Europe PMC Provider (
src/linkml_reference_validator/etl/fulltext/epmc.py):FullTextProviderRegistryas "epmc"Updated Provider Registry (
src/linkml_reference_validator/etl/fulltext/__init__.py):Comprehensive Tests (
tests/test_fulltext_providers.py):Implementation Details
config.rate_limit_delayMIN_FULLTEXT_CHARS) to return resultshttps://claude.ai/code/session_01PYbfu1JPi9rr7v5XU6qbJj