pyEuropePMC

PyEuropePMC Features

Explore what PyEuropePMC can do - Comprehensive feature overview and workflows

Core Features

Query the Europe PMC database with powerful search capabilities

Quick Example:

from pyeuropepmc import SearchClient

with SearchClient() as client:
    results = client.search("cancer AND therapy", pageSize=50, sort="CITED desc")

Learn More

Full-Text Retrieval

Download complete article content in multiple formats

Quick Example:

from pyeuropepmc import FullTextClient

with FullTextClient() as client:
    pdf_path = client.download_pdf_by_pmcid("PMC1234567")
    xml_content = client.download_xml_by_pmcid("PMC1234567")

Learn More

XML Parsing

Extract structured data from full-text XML documents

Quick Example:

from pyeuropepmc import FullTextXMLParser, ElementPatterns

parser = FullTextXMLParser(xml_content)

# Extract metadata
metadata = parser.extract_metadata()

# Extract tables
tables = parser.extract_tables()

# Convert to markdown
markdown = parser.to_markdown()

# Validate schema coverage
coverage = parser.validate_schema_coverage()
print(f"Coverage: {coverage['coverage_percentage']:.1f}%")

Learn More

Query Builder

Advanced fluent API for building complex search queries with type safety

Quick Example:

from pyeuropepmc import QueryBuilder

qb = QueryBuilder()
query = (qb
    .keyword("cancer", field="title")
    .and_()
    .citation_count(min_count=50)
    .and_()
    .date_range(start_year=2020)
    .build())
# Result: "(TITLE:cancer) AND (CITED:[50 TO *]) AND (PUB_YEAR:[2020 TO *])"

Learn More

Systematic Review Tracking

PRISMA/Cochrane-compliant search logging and audit trails

Quick Example:

from pyeuropepmc import QueryBuilder
from pyeuropepmc.utils.search_logging import start_search

log = start_search("Cancer Review", executed_by="Researcher")
qb = QueryBuilder().keyword("cancer").and_().field("open_access", True)
qb.log_to_search(log, filters={"open_access": True}, results_returned=100)

Learn More

Feature Comparison

Feature SearchClient FullTextClient FullTextXMLParser FTPDownloader QueryBuilder
Search Europe PMC Yes - - - Yes
Build Complex Queries - - - - Yes
Type-Safe Fields - - - - Yes
Query Validation - - - - Yes
Query Translation - - - - Yes
Download PDFs - Yes - Yes -
Download XML - Yes - - -
Parse XML - - Yes - -
Extract Metadata - - Yes - -
Extract Tables - - Yes - -
Bulk Downloads - - - Yes -
Systematic Review Logging - - - - Yes
Caching Yes Yes - - -
Progress Tracking - Yes - Yes -

Common Workflows

Workflow 1: Advanced Query -> Search -> Parse

from pyeuropepmc import QueryBuilder, SearchClient, FullTextXMLParser

# Step 1: Build complex query with QueryBuilder
qb = QueryBuilder()
query = (qb
    .keyword("machine learning", field="title")
    .and_()
    .citation_count(min_count=25)
    .and_()
    .date_range(start_year=2020)
    .build())

# Step 2: Search with the query
with SearchClient() as client:
    results = client.search(query, pageSize=20, sort="CITED desc")

    # Step 3: Process results
    for paper in results['resultList']['result']:
        if paper.get('pmcid'):
            # Download and parse XML
            xml_content = client.get_fulltext_xml(paper['pmcid'])
            parser = FullTextXMLParser(xml_content)
            metadata = parser.extract_metadata()
            print(f"High-impact paper: {metadata['title']}")

Workflow 2: Systematic Review with Audit Trail

from pyeuropepmc import QueryBuilder
from pyeuropepmc.utils.search_logging import start_search

# Start systematic review
log = start_search("ML in Biology Review", executed_by="Researcher Name")

# Build comprehensive search strategy
qb = QueryBuilder()
comprehensive_query = (qb
    .keyword("machine learning")
    .and_()
    .keyword("biology")
    .and_()
    .field("open_access", True)
    .and_()
    .date_range(start_year=2019)
    .build())

# Execute and log search
with SearchClient() as client:
    results = client.search(comprehensive_query, pageSize=100)

    # Log for systematic review compliance
    qb.log_to_search(
        search_log=log,
        filters={"open_access": True, "date_range": "2019+"},
        results_returned=len(results['resultList']['result']),
        notes="Comprehensive ML in biology search"
    )

# Save review log
log.save("systematic_review_log.json")

Workflow 3: Advanced Search -> Filter -> Extract

from pyeuropepmc import SearchClient, FullTextXMLParser

with SearchClient() as client:
    # Advanced search with filters
    results = client.search(
        query="cancer AND (therapy OR treatment)",
        sort="CITED desc",
        pageSize=100,
        resultType="core"
    )

    # Filter for high-impact papers
    high_impact = [
        paper for paper in results['resultList']['result']
        if paper.get('citedByCount', 0) > 50 and paper.get('pmcid')
    ]

    # Extract detailed information
    for paper in high_impact:
        xml_content = client.get_fulltext_xml(paper['pmcid'])
        parser = FullTextXMLParser(xml_content)
        # Analyze...

Feature Matrix

Search Features

Capability Supported Notes
Keyword search Yes Full-text search across all fields
Boolean operators Yes AND, OR, NOT
Field-specific Yes Search specific fields (author, title, etc.)
Date filtering Yes Publication date ranges
Citation sorting Yes Sort by citation count
Pagination Yes Handle large result sets
Multiple formats Yes JSON, XML, Dublin Core

Full-Text Features

Capability Supported Notes
PDF download Yes Open access articles only
XML download Yes JATS/NLM XML format
HTML content Yes HTML representation
Bulk FTP Yes Efficient for large datasets
Progress tracking Yes Real-time progress callbacks
Auto-retry Yes Robust error handling

Parsing Features

Capability Supported Notes
Metadata extraction Yes Title, authors, journal, dates, etc.
Table extraction Yes Structured table data
Reference extraction Yes Complete bibliography
Plaintext conversion Yes Full article text
Markdown conversion Yes Formatted markdown
Schema validation Yes Coverage analysis
Custom patterns Yes Flexible configuration
Multiple XML schemas Yes JATS, NLM, custom

Learning Resources

By Feature

By Use Case

By Skill Level

Best Practices

Performance

Error Handling

Rate Limiting

Section Why Visit?
Getting Started Installation and basics
API Reference Complete method documentation
Examples Working code samples
Advanced Power user features