AEGIS2.0: a diverse AI safety dataset and risks taxonomy
I co-authored AEGIS2.0, a diverse AI safety dataset and risks taxonomy for alignment of LLM guardrails.
The work introduces a scalable taxonomy of 12 core hazard categories and 9 fine-grained risks, plus a hybrid data pipeline that combines human annotation with a multi-LLM jury. The resulting dataset of human–LLM interactions is designed for commercial use, and models trained on it are competitive with leading safety systems trained on much larger, non-commercial data.