Dataset for "Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis"

Source list used to build the specialised bodies of text we used in our research.

This excel dataset includes a complete list of sources used to compile the corpora used in data analysis. Fields include publisher, heading, file type, and URL or/and or other source origin information where relevant.

Keywords:
Corporate communication, Corporate language, Tobacco, health information
Subjects:

Cite this dataset as:
Fitzpatrick, I., Sun, X., 2026. Dataset for "Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis". Bath: University of Bath Research Data Archive. Available from: https://doi.org/10.15125/BATH-01596.

Export

Data

List of Data … HOC and TIC.xlsx
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet (117kB)
Creative Commons: Attribution 4.0

Complete list of sources used to compile corpora used in the analysis of our paper. Please note that for archived tobacco industry webpages in cases where the original publication date wasn’t available, the date last updated was used; if that wasn’t available either, the date archived was used.

This dataset contains data extracted from corporate websites. The Information was available in the public domain at the time of data gathering. We cannot guarantee continued public access to the data.

Creators

Xinmei Sun
Lancaster University

Contributors

University of Bath
Rights Holder

Coverage

Temporal coverage:

From 1 January 2003 to 31 December 2023

Documentation

Data collection method:

The search parameters for the tobacco industry corpus (TIC) were: web pages and documents published by British American Tobacco (BAT) and Philip Morris International (PMI). We used the following search query to capture contents containing terms related to smoking status and smoking behaviour: smok* OR quit* OR cessation OR switch* OR vap* OR experiment* OR relaps* OR "us* NRT*" OR "dual" OR addict*. We carried out three separate searches: 1) Google site search, 2) using the search function within the two tobacco company websites, and 3) searching web archive tools – The Internet Archive and archive.today – to collect back dated content not captured in searches 1 or 2.

Data processing and preparation activities:

The list of returned results were saved as structured CSV files containing metadata such as URLs, headings, publication dates. This list was then manually filtered to exclude job adverts, pages spotlighting named individuals, and pages or documents in which search terms only appeared in cautionary statements or legal disclosures. We also excluded region-specific/geographically local content pages.

Technical details and requirements:

We used Python scripts to process the CSV files. The scripts included routines for downloading linked PDF documents, extracting and cleaning text from PDF documents using the pdfminer library (https://pypi.org/project/pdfminer/), and exporting web-based content as individual plain text (.txt) files.

Funders

Bloomberg Philanthropies
https://doi.org/10.13039/100015283

STOP 2

Publication details

Publication date: 15 July 2026
by: University of Bath

Version: 1

DOI: https://doi.org/10.15125/BATH-01596

URL for this record: https://researchdata.bath.ac.uk/1596

Related papers and books

Fitzpatrick, I., and Sun, X., 2026. Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis. JMIR Infodemiology, 6, e86593-e86593. Available from: https://doi.org/10.2196/86593.

Contact information

Please contact the Research Data Service in the first instance for all matters concerning this item.

Contact person: Iona Fitzpatrick

Departments:

Faculty of Humanities & Social Sciences
Health