FSCI 2026 | E01_AI-Assisted Literature Review

The FSCI2026 is an international, community-driven training program focused on advancing modern scholarly communication practices. It brings together researchers, librarians, developers, and open science advocates to explore topics like open access, research data management, reproducibility, and AI in research.

Read more

Theme

Critical Thinking and Evolving Open Scholarship in the Era of AI

semanticClimate team is excited to participate as course instructors at the FSCI2026.

Dates

Mode

Course Information

Read the Course abstract

E01: AI-assisted Literature Review: Building Open, Machine-Readable Literacy for Research and Education, Featuring the Semantic Climate Encyclopedia

Date : 27–31 July 2026

Class Schedule

Instructors

Course chair: Simon Worthington, semanticClimate, Publishing Knowledge Graph Researcher (TIB – Leibniz Information Centre for Science and Technology and University Library) Dr. Gitanjali Yadav, BRIC-National Institute of Plant Genome Research (BRIC-NIPGR) (Co-course chair); Prof. Peter Murray-Rust, Cambridge University (Co-course chair); Dr. Renu Kumari, #semanticClimate (Instructor), Ms. Priti, BRIC-National Institute of Plant Genome Research (Instructor), Ms. Udita Agarwal, BRIC-National Institute of Plant Genome Research (Instructor)

Abstract

The exponential growth of scholarly literature across disciplines has created major challenges for discovering, synthesizing, and reusing knowledge in a systematic and scalable manner. Traditional literature review practices and static reference resources are increasingly inadequate for handling large, heterogeneous, and rapidly evolving information landscapes. The course demonstrates how AI-assisted literature review workflows can be built, evaluated, and adapted using open access scholarly literature and transparent, reproducible methods.

The proposed framework leverages semantic technologies, knowledge graphs, and natural language processing to transform unstructured publications, reports, and documents into structured, interoperable, and reusable knowledge representations. AI-assisted literature review supports automated, transparent, and continuously updatable synthesis of evidence, enabling efficient navigation of large corpora and reducing manual effort. Alignment with shared vocabularies and standards ensures consistency, interoperability, and long-term reusability of the curated knowledge.

A core component of the framework is Named Entity Recognition (NER), which enables the automated identification and classification of key entities—such as species name, organizations, locations, diseases, drugs from research publications. NER provides the foundation for linking entities across documents, establishing semantic relationships, and populating knowledge graphs that support advanced querying, comparison, and reasoning.

Central to this effort is the Semantic Climate Encyclopedia—a dynamic, semantic reference that systematically enriches keywords with linked data from Wikimedia sources, and organizes them into a structured, searchable knowledge system. Semantic Encyclopedias can be automatically created for any corpus and are an excellent way for researchers to get an introduction to a new subject within an hour or two.

Together, the Semantic Encyclopedia, AI-assisted review, Named Entity Recognition (NER), and LLM-powered Retrieval-Augmented Generation (RAG) form a dynamic, extensible infrastructure that improves research efficiency, enhances knowledge accessibility, and supports data-driven learning and decision-making across disciplines. This integrated approach advances open science by bridging human-readable scholarship with machine-interpretable knowledge, enabling scalable, reusable, and interoperable research and educational ecosystems.

Course units include:

The course will enable to build, deploy, and adapt an AI-assisted literature review workflow for semi-automated discovery, retrieval, and synthesis of peer-reviewed research articles. The instructional example uses scientific chapters and references from large open access article collections comprising millions of research articles. Participants will be introduced to core AI and machine learning concepts for literature analysis, including document retrieval, text mining, NER, and LLM-based retrieval systems, with a strong emphasis on evaluation, bias awareness, and the trustworthiness of AI-assisted outputs.

Audience:

Researchers Researchers, librarians, publishers

Level: (Beginner, but suitable for all levels)

Requirements:

Run Google Colab in a browser. See Google Colab A Google account to run CoLab Jupyter Notebooks

Course Learning Objectives

At the end of the course, participants will be able to:

Course Topics

This course will be presented over three days for 2 hours each day and will cover these topics:

Day 1, Tuesday, July 28, 2026

Topic 1: Knowledge extraction from scientific literature corpus

Software/Tools used:

Description:

Knowledge extraction from a scientific literature corpus involves transforming large volumes of unstructured research articles into structured, machine-readable knowledge that can be queried, analyzed, and reused. Using the #semanticClimate toolkits particularly pygetpapers, amilib, and docanalysis, this process becomes automated, scalable, and reproducible.

The workflow begins with pygetpapers, which retrieves relevant open-access scientific articles in bulk using keyword-based queries. These articles are downloaded in structured formats such as XML, enabling downstream computational analysis. The collected corpus is then visualized as datatable using amilib. The datatable is browsable and searchable. This contains information about title, author, year of publication, DOIs etc. for all the retrieved articles.

Once a clean and structured corpus is prepared, docanalysis is used to perform detailed text mining and natural language processing. It enables section-wise parsing, dictionary-based annotation, and named entity recognition (NER) to extract key information such as organisms, chemicals, locations, diseases, or research concepts. The tool can also generate domain-specific dictionaries and structured outputs (CSV/JSON), facilitating downstream analysis and knowledge graph construction.

Together, these tools form an integrated pipeline that supports automated literature review, keyword extraction, and semantic enrichment. The extracted entities and relationships can be organized into knowledge graphs or encyclopedic resources, enabling deeper insights, pattern discovery, and data-driven research across scientific domains. This approach significantly enhances the efficiency, reproducibility, and scalability of knowledge extraction from large scientific corpora.

Day 2, Wednesday, July 29, 2026

Topic 2: Create semantic encyclopedia for knowledge enrichment

Software/Tools used:

Description:

Creating a semantic encyclopedia for knowledge enrichment involves organizing extracted concepts and terms from scientific literature into a structured, searchable, and interconnected resource. Using the #semanticClimate tools txt2phrases and encyclopedia, this process enables the transformation of raw textual data into a domain-specific knowledge base that supports exploration, discovery, and reuse.

The workflow involves txt2phrases, which analyzes cleaned text corpora to automatically identify meaningful keyphrases from the research articles. This helps in generating a rich vocabulary that reflects how knowledge is actually expressed within a specific research domain. These extracted keyphrases are then passed to the encyclopedia tool, which structures them into an organized semantic resource. The encyclopedia tool allows users to create entries, define relationships between concepts, and enrich terms with contextual information such as definitions, figures from the trustable resource Wikimedia. The result is a machine-readable and human-interpretable knowledge system that can function as a lightweight domain ontology or glossary.

Together, txt2phrases and encyclopedia enable the creation of a continuously evolving semantic encyclopedia that enhances knowledge discovery, supports automated reasoning, and facilitates integration with knowledge graphs and FAIR data systems.

Day 3, Thursday, July 30, 2026

Topic 3: LLM and RAG for processing and querying scientific documents

Description:

This course explains how Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) can be used to work with PDF and HTML documents in a simple and efficient way. Participants will learn how to extract useful information from large documents, ask questions, and get accurate answers based on the actual content.

It also introduces how to break documents into smaller sections, store them for quick search, and connect them with LLMs to improve responses. The module focuses on practical, step-by-step workflows that help in reading, summarizing, and exploring large collections of text.

Participants will gain hands-on experience in building simple systems that can search across multiple documents and provide reliable answers. The unit also highlights ways to reduce errors, improve accuracy, and make better use of digital documents for research and learning.

Co-course Chair:

Mr. Simon Worthington

Mr. Simon Worthington
TIB – Leibniz Information Centre for Science and Technology and University Library, Germany; #semanticClimate

Dr. GY

Dr. Gitanjali Yadav
#semanticClimate; BRIC-National Institute of Plant Genome Research (BRIC-NIPGR), New Delhi, India

PMR

Prof. Peter Murray-Rust
#semanticClimate; Cambridge University, UK

Course Speaker and Instructors

Dr. Renu Kumari

Dr. Renu Kumari
#semanticClimate; BRIC-NIPGR, India

Ms. Priti

Ms. Priti Chahal
#semanticClimate; BRIC-NIPGR, India

Ms. Udita Agarwal

Ms. Udita Agarwal
#semanticClimate; BRIC-NIPGR, India

Program Schedule and course instructors

Date and Time Title Instructor
Tuesday, 28 July, 2026 Knowledge extraction from scientific literature corpus Dr. Renu Kumari
Wednesday, 29 July, 2026 Create semantic encyclopedia for knowledge enrichment Ms. Udita Agarwal
Thursday, 30 July, 2026 LLM and RAG for processing and querying scientific documents Ms. Priti Chahal

FSCI Opening Plenary Session, July 27, 2026

Keynote : "The Importance of the ENVISAGE Principles and PILOT Principles for AI Governance in Science Communications" by Dr Jacintha Toohey and Francis P. Crawley.

Recording of the session

Day 1 Session, July 28, 2026, Knowledge extraction from scientific literature corpus

Recording of the session

Day 2 Session, July 29, 2026, Create Semantic Encyclopedia for Knowledge Enrichment

Recording of the session

Day 3 Session, July 30, 2026, LLM and RAG for processing and querying scientific documents

Recording of the session

Closing Plenary Panel

Title : AI: Pedagogical Partner or Scholarly Disruptor? Read Abstract

Friday July 31, 2026

8:00 AM – 10:00 AM Pacific time (UCT-7)

Recording of the session

Speakers of Closing Plenary Panel

speaker-1

Shana Pote is a doctoral candidate, educator, and executive coach whose work focuses on how emerging technologies reshape human judgment, authorship, and knowledge production. Read more

speaker-2

Nassima Berrazouane is a Ph.D. student at the University of Tiaret, Algeria. Read more

speaker-3

Hande Kucuk McGinty, Ph.D. is an Assistant Professor in the Department of Computer Science at Kansas State University. Read more

speaker-4

Aaron Tay is an academic librarian. He currently leads research and data services at Singapore Management University Libraries. Read more

Closing Plenary B

Title : When We Said “Open,” We Didn’t Mean “OpenAI”: OA, Megacorps, and the Illusion of Control Read Abstract

Friday July 31, 2026

11:00 AM – 1:00 PM Pacific time (UCT-7)

Recording of the session

Speaker

Speaker

Dave Hansen is the Executive Director of Authors Alliance. Read more

Organizations We Thank

BRIC-NIPGR

BRIC- National Institute of Plant Genome Research (BRIC-NIPGR), India

CODATA

CODATA (Committee on Data of the International Science Council)

FSCI 2026

FSCI 2026

FORCE11

FORCE11 (The Future of Research Communications and e-Scholarship)

semanticClimate

semanticClimate

event semanticclimate outreach hackathon

← Back