Dataset for "Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis"
Source list used to build the specialised bodies of text we used in our research.
This excel dataset includes a complete list of sources used to compile the corpora used in data analysis. Fields include publisher, heading, file type, and URL or/and or other source origin information where relevant.
Cite this dataset as:
Fitzpatrick, I.,
Sun, X.,
2026.
Dataset for "Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis".
Bath: University of Bath Research Data Archive.
Available from: https://doi.org/10.15125/BATH-01596.
Export
Data
List of Data … HOC and TIC.xlsx
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet (117kB)
Creative Commons: Attribution 4.0
Complete list of sources used to compile corpora used in the analysis of our paper. Please note that for archived tobacco industry webpages in cases where the original publication date wasn’t available, the date last updated was used; if that wasn’t available either, the date archived was used.
This dataset contains data extracted from corporate websites. The Information was available in the public domain at the time of data gathering. We cannot guarantee continued public access to the data.
Contributors
University of Bath
Rights Holder
Coverage
Temporal coverage:
From 1 January 2003 to 31 December 2023
Documentation
Data collection method:
The search parameters for the tobacco industry corpus (TIC) were: web pages and documents published by British American Tobacco (BAT) and Philip Morris International (PMI). We used the following search query to capture contents containing terms related to smoking status and smoking behaviour: smok* OR quit* OR cessation OR switch* OR vap* OR experiment* OR relaps* OR "us* NRT*" OR "dual" OR addict*. We carried out three separate searches: 1) Google site search, 2) using the search function within the two tobacco company websites, and 3) searching web archive tools – The Internet Archive and archive.today – to collect back dated content not captured in searches 1 or 2.
Data processing and preparation activities:
The list of returned results were saved as structured CSV files containing metadata such as URLs, headings, publication dates. This list was then manually filtered to exclude job adverts, pages spotlighting named individuals, and pages or documents in which search terms only appeared in cautionary statements or legal disclosures. We also excluded region-specific/geographically local content pages.
Technical details and requirements:
We used Python scripts to process the CSV files. The scripts included routines for downloading linked PDF documents, extracting and cleaning text from PDF documents using the pdfminer library (https://pypi.org/project/pdfminer/), and exporting web-based content as individual plain text (.txt) files.
Funders
Bloomberg Philanthropies
https://doi.org/10.13039/100015283
STOP 2
Publication details
Publication date: 15 July 2026
by: University of Bath
Version: 1
DOI: https://doi.org/10.15125/BATH-01596
URL for this record: https://researchdata.bath.ac.uk/1596
Related papers and books
Fitzpatrick, I., and Sun, X., 2026. Investigation of Differences in Tobacco Use Language Between Groups: Corpus-Assisted Analysis. JMIR Infodemiology, 6, e86593-e86593. Available from: https://doi.org/10.2196/86593.
Contact information
Please contact the Research Data Service in the first instance for all matters concerning this item.
Contact person: Iona Fitzpatrick
Faculty of Humanities & Social Sciences
Health