# Mining experimental data from materials science literature with large language models: an evaluation study

https://mdr.nims.go.jp/datasets/440e1b12-8985-4302-88f0-c880a8116aab

## File

- [Mining experimental data from materials science literature with large language models  an evaluation study.pdf](https://mdr.nims.go.jp/filesets/29761a34-e572-4da0-aa29-48ce0fcdfaa0/download) ([Detail](https://mdr.nims.go.jp/filesets/29761a34-e572-4da0-aa29-48ce0fcdfaa0.md))

## Id

440e1b12-8985-4302-88f0-c880a8116aab

## Local identifier



## Visibility

open_to_public

## State

published

## Created at

2024-10-23T07:55:18.116575Z

## Updated at

2024-10-24T07:30:23.218419Z

## Published at

2024-10-24T07:30:24.853965Z

## Doi



## First published url

https://doi.org/10.1080/27660400.2024.2356506

## Date published

2024-12-31

## Recorded date published

2024-12-31

## Resource type

journal_article

## Manuscript type

vor

## Collection



## Title

- title: 'Mining experimental data from materials science literature with large language
    models: an evaluation study'
  title_type: original
  lang: en

## Description

- description: This study is dedicated to assessing the capabilities of large language
    models (LLMs) such as GPT-3.5-Turbo, GPT-4, and GPT-4-Turbo in the extraction
    of structured information from scientific documents in materials science. To this
    end, we primarily focus on (i) a named entity recognition (NER) of studied materials
    and physical properties and (ii) a relation extraction (RE) between these entities.
    The performance of LLMs in executing these tasks is benchmarked against traditional
    models, BERT and rule-based approaches. As a typical result, GPT-4 and GPT-4-Turbo
    display remarkable reasoning and relationship extraction capabilities after being
    provided with merely a couple of examples.
  description_type: abstract
  lang: und

## Creator

- name: Luca Foppiano
  role: author
  orcid: https://orcid.org/0000-0002-6114-6164
- name: Guillaume Lambard
  role: author
  orcid: https://orcid.org/0000-0003-0275-4079
- name: Toshiyuki Amagasa
  role: author
- name: Masashi Ishii
  role: author
  orcid: https://orcid.org/0000-0003-0357-2832

## Contact agent



## Publisher

organization: Informa UK Limited

## Managing organization



## Keyword

- subject: Large language models
  schema: not_defined
- subject: benchmark
  schema: not_defined
- subject: NER
  schema: not_defined
- subject: TDM
  schema: not_defined
- subject: evaluation
  schema: not_defined
- subject: materials science
  schema: not_defined

## Rights

- identifier: https://creativecommons.org/licenses/by/4.0/

## Other identifier(s)



## Data origin



## Embargo



## Journal

- title: 'Science and Technology of Advanced Materials: Methods'
  issn: '27660400'
  volume: '4'
  issue: '1'

## Conference



## Related item



## Funding

- identifier: JPMXP1122715503
  funder_name: Research and Development

## Instrument



## Instrument operator



## Instrument managing organization



## Measurement method



## Specimen



## Chemical composition



## Structure for specimen



## Structural feature for specimen



## Specific property for specimen



## Process for specimen treatment



## Computational method



## Energy level/transition state



## Software



## Custom property



## Fileset

- id: 29761a34-e572-4da0-aa29-48ce0fcdfaa0
  filename: Mining experimental data from materials science literature with large
    language models  an evaluation study.pdf
  content_type: application/pdf
  size: 3346772
  md5: a875c11dff21f0c4bf4c01b8c128dab1

## Thumbnail

fileset_id: 29761a34-e572-4da0-aa29-48ce0fcdfaa0
filename: Mining experimental data from materials science literature with large language
  models  an evaluation study.pdf