Sources and data policy

What CSI uses, and what it does not use

The MVP is open-data only. Each public profile should show where the event, speaker, author-metric, baseline, and score inputs came from.

Open-data policy

CSI uses public APIs, official conference pages, society calendars, and manually approved source snapshots. The service stores raw responses or source captures before normalization so later reviewers can audit what changed.

Source availability varies by field. Computer science has stronger structured venue data through DBLP and similar indexes. Chemistry, materials science, and polymer science rely more on official sites, society calendars, and manual curation.

MVP source classes

SourceUsed forPolicy
OpenAlexAuthor metrics and source-level citation-intensity data for the Field Citation Baseline.Automated API use is allowed with responsible rate limits and configured contact details.
Semantic ScholarAuthor search and public author metric support where records can be matched with enough confidence.Use public or keyed API access. Store raw responses before normalization.
DBLPVenue discovery and conference metadata, mainly for computer science fields.Use documented search/export endpoints and keep normalized records traceable to source responses.
Official conference sitesDates, locations, speaker lists, roles, program pages, official URLs, and proceedings links.Capture source snapshots and label organizer-supplied fields where official-site data is used.
Society calendarsRecurring event discovery and schedule updates from scholarly societies and publishers.Use calendars or exports only where terms and access allow recurring ingestion.

Google Scholar

Manual supporting evidence only; do not automate scraping or bulk extraction.

Public Google Scholar profile links or manually captured h-index values may appear as supporting evidence when a reviewer records them. CSI should not scrape Google Scholar search results or profiles, and it should not bulk-extract h-index values from Google Scholar.

Paid or restricted data

The open-data MVP does not use licensed Journal Impact Factor data, Scopus metrics, or other paid bibliometric feeds for the public score.

  • Clarivate Journal Citation Reports and Web of Science
  • Scopus author and source APIs
  • Dimensions or other licensed bibliometric datasets
  • Google Scholar automated scraping

Traceability rules

Known limitations

Coverage will be uneven at launch. Some conferences publish clear speaker lists and program pages. Others publish partial schedules, image-only PDFs, or delayed speaker pages. Low coverage and low confidence labels are there to show that limitation instead of hiding it.

CSI is a bibliometric signal based on verified plenary and keynote speakers. It is not a measure of community value, mentoring quality, acceptance rigor, attendee experience, or whether a conference is the right venue for a specific paper.