Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

2.1 Overview 

  •  Create a new diagram to have clearly visible texts
  •  link these subheadings also to main subsections with the nos of subchapters in this manual
  •  Add more references
  •  

The general data workflow outlines how bioarchaeological biomolecular and archaeological data is are managed from its their generation through various usage and up to structured archival. The diagram below visualizes Figure 1#below visualises the core progression, showing data infrastructure, in this case based on SharePoint worksheets and ARHUT data management system and emphasizing emphasising feedback loops in data management: 

Image Added

Figure 1. Image Removed

2.2 Sampling and Initial Documentation 

Samplingisbased on researchdesign and must followconsistentdocumentationpractices. Eachsampleisgiven a uniqueidentifier following local laboratory principles of sample labelling, withmetadatacoveringobject/artefact type,itscollection number, excavationcontext, coordinates, objecttype, sample collectiondate, sampler, and analysismethodetc planned.Documentationbegins in fieldorlabnotebooks-books and islatertranscribedinto SharePoint worksheets.cloud-based worksheets (e.g., SharePoint, Google Docs). 

References

Niven, K., Jakobsson, U. Databases and spreadsheets: A guide to good practicehttps://zenodo.org/records/7740647  

MINAS: (DNA)http://www.mixs-minas.org/ 

Stable isotopes: https://doi.org/10.1016/j.quaint.2022.02.027;  

Roberts P, Fernandes R, Craig OE, Larsen T, Lucquin A, Swift J, Zech J. Calling all archaeologists: guidelines for terminology, methodology, data handling, and reporting when undertaking and reviewing stable isotope applications in archaeology. Rapid Commun Mass Spectrom. 2018 Mar 15;32(5):361-372. doi: 10.1002/rcm.8044. PMID: 29235694; PMCID: PMC5838555. 

 

Reiter, Samantha S., Staniuk, Robert, Kolář, Jan, Bulatović, Jelena, Rose, Helene Agerskov, Ryabogina, Natalia E., Speciale, Claudia, Schjerven, Nicoline, Paulsson, Bettina Schulz, Lee, Victor YanKin, Canteri, Elisabetta, Revill, Alice, Dahlberg, Fredrik, Sabatini, Serena, Frei, Karin M., Racimo, Fernando, Ivanova-Bieg, Maria, Traylor, Wolfgang, Kate, Emily J., Derenne, Eve, Frank, Lea, Woodbridge, Jessie, Fyfe, Ralph, Shennan, Stephen, Kristiansen, Kristian, Thomas, Mark G. and Timpson, Adrian. "The BIAD Standards: RecommendationsforArchaeologicalDataPublication and InsightsFromtheBigInterdisciplinaryArchaeologicalDatabase" OpenArchaeology, vol. 10, no. 1, 2024, pp. 20240015. https://doi.org/10.1515/opar-2024-0015 

2.3 DataAcquisition and Initial

...

Recording 

Instrument outputssuch as LC-MSoutputs–such as mass-spectrometry (IRMS, GC-MS, NGS LC-MS/MS) andsequencing files, ormicroscopydata—are worksheetsvisualsare collected in vendor-specificrawformats (e.g., RAW, FASTQ), and preferably stored in instrument-related computer and copied into project (shared) folders, securing the back-up versions of initial measurement files. Thisrawdataisthenreferenced and linked in SharePoint combined worksheets (e.g. SharePoint or Google Sheets)thatrecordinitialmetadata, samplingcontext, and lab-specificidentifiers. Theseworksheets are usedforearly-stagereview and validation. 

2.4 DataStructuring and CollaborativeEditing 

Rawentries are transformedintostructuredresearchdatasets by moving them to tabular data sheets (e.g. Excel), cleaning data, standardisingterminology, and checkingforconsistency. Thisstepincludes: 

  • Keeping consistent data structures by harmonising column names and formats to keep them consistent within the work process. 
    Selecting relevant data fields, harmonising column names and formats
  • Selecting relevant data fields that will be filled/edited during the given datasets/stages
  •  
  • Validating entries against the requirement in the original documentation, e.g. field formats, required label, etc. 
  • Assigning
  • internal project
  • relational identifiers (e.g. site code, ledger number) and internal project/lab codes and versioning identifiersinto corresponding fields of data sheets. 

These structured datasets form the basis for computational analysis and are maintained within SharePoint for collaborative editing (e.g. “live” editing for Excel, but for other files, it might include different versions edited by different people). Access permissions are set to control changes and ensure data provenance. 

2.5 Data Analysis 

Once structured, datasets can be exported (typically as Comma-Separated Values, CSV files) and processed using computational tools tailored to specific research questions, analyses and data types. This includes data interpretation and evaluation, statistical modelling, pattern recognition, and visualisation. Analyses are typically performed in environments like R, Python, or specialised software, such as OxCal, IsoReader, mMass, or MaxQuant. 

Analytical outputs must be reproducible and versioned, with all scripts and parameter settings documented and stored alongside the dataset, either in SharePoint or linked repositories (e.g., GitHub). 

 

2.6 Data Validation and Feedback 

Structured datasets are subjected to both planned and unplanned quality checks. Users can verify data completeness, coherence, and consistency with raw entries during a formal review process, but often various problems arenoticed while working with data. In some cases, those require contextual knowledge, and it is thus not possible to catch all of those during any formal review. Feedback is communicated via SharePoint comments , or tracked changes, or or ARHUT comments / tasks. Datasets in case the dataset has already been entered to the ARHUT system. Datasets may cycle through multiple revisions before finalisationfinalization.This feedback mechanism is essential for maintaining data quality and for correcting inconsistencies before deposition. 

2.7 Curation and Archival in ARHUT 

FinalisedFinalizeddatasets are transferredtothethe ARHUT data platform, wherethey are archivedwith: 

  • Persistent
  • identifiers 
  • identifiers. Each entity has its own ARHUT link, that can be used to reference from publications but is essential in linking datasets. Additionally other identifiers can be added e.g.dataDOI.  
  • Full metadata including sampling context, lab identifiers, and data structure 
  • Relations to other data tables
  • within the
  • within the system, forming agradually densifying knowledge graph. 

Thesedatasetsbecome part of thelong-term record and are linkedtobothinternalsystems (e.g., SharePoint, Archemy, Department of Archaeology) and externalrepositories (e.g.  Zenodo,Dryad, BIAD ,etc Dryad) and databases (e.gBIAD). 

2.8 Storage Platforms and File Formats 

  •  Add another diagram here about the structure of what is happening

Eachphase of theworkflowissupported by designatedplatforms: 

  • SharePoint for live collaboration and version-controlled documentation 
  • ARHUT
  • for
  •  for curated, long-term data with controlled access and open publishing options 
  • Lab databases (e.g.  BBAD) for supplementary metadata and internal tracking 

2.9 Dissemination 

Finalised and curated datasets archived in in ARHUT are made available through their dissemination. ARHUT's web interface (https://arh.ut.ee/) allows for structured querying and access to project-specific datasets, enriched with contextual metadata and persistent identifiers. Paleomix Open Archaeological Database (PaleoMIX O.A.D.) builds on the the ARHUT infrastructure, offering offering public-facing access to selected datasets from Paleomix PaleoMIX and related projects. This system enables transparent sharing of research outputs, supports interdisciplinary collaboration, and fosters broader reuse by both academic and public audiences.  


References

Reiter, Samantha S., Staniuk, Robert, Kolář, Jan, Bulatović, Jelena, Rose, Helene Agerskov, Ryabogina, Natalia E., Speciale, Claudia, Schjerven, Nicoline, Paulsson, Bettina Schulz, Lee, Victor Yan Kin, Canteri, Elisabetta, Revill, Alice, Dahlberg, Fredrik, Sabatini, Serena, Frei, Karin M., Racimo, Fernando, Ivanova-Bieg, Maria, Traylor, Wolfgang, Kate, Emily J., Derenne, Eve, Frank, Lea, Woodbridge, Jessie, Fyfe, Ralph, Shennan, Stephen, Kristiansen, Kristian, Thomas, Mark G. and Timpson, Adrian. "The BIAD Standards: Recommendations for Archaeological Data Publication and Insights From the Big Interdisciplinary Archaeological Database" Open Archaeology, vol. 10, no. 1, 2024, pp. 20240015. https://doi.org/10.1515/opar-2024-0015