Visible AI assistant outputs in psychosocial risk management: an OHP/OHS-grounded benchmark for governance and responsible use


Creative Commons License

ÇÖGENLİ M. Z.

Frontiers in public health, cilt.14, ss.1857113, 2026 (SCI-Expanded, SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 14
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3389/fpubh.2026.1857113
  • Dergi Adı: Frontiers in public health
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Social Sciences Citation Index (SSCI), Scopus, EMBASE, MEDLINE, Psycinfo, Directory of Open Access Journals
  • Sayfa Sayıları: ss.1857113
  • Anahtar Kelimeler: AI assistants, benchmarking, governance, ISO 45003, occupational health and safety, occupational health psychology, psychosocial risk management, reporting
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Uşak Üniversitesi Adresli: Evet

Özet

Background: This study introduces and applies an occupational health psychology (OHP) and occupational health and safety (OHS)-grounded benchmark for evaluating visible artificial intelligence (AI) assistant outputs in psychosocial risk management. Methods: Five user-facing general-purpose AI assistant products were evaluated across four locked psychosocial-risk scenarios, three repeated runs, and four fixed tasks per conversation, yielding 60 conversations and 240 task-level outputs. Outputs were scored on risk identification and differentiation, multi-level organizational framing, preventive organizational actionability, and professional boundedness and verification. A blind second-rater layer and focused adjudication were used to test profile stability and resolve profile-changing disagreements. Results: Exact agreement was observed in 153 of 240 task rows (63.7%), while 239 rows (99.6%) were exact or within one point. Profile mismatches were concentrated in Gemini and Le Chat/Mistral rather than distributed across scenarios or runs. Under the final consensus structure, 36 conversations were classified as Structurally Usable Support, 23 as Mixed/Partially Usable Support, and 1 as Problematic Support. Conclusion: The main contribution is not a generic ranking of assistant products, but a domain-specific evaluation logic for psychosocial risk management, together with a reviewer-facing process model and derived governance and reporting guides for more transparent and responsible organizational use.