
Elith Inc. announces that it participated in "Open Safeguard Hackathon," an international hackathon co-hosted by ROOST, Hugging Face, and OpenAI, held in San Francisco, USA on December 8, 2025.

The hackathon was held as a venue to practically verify and develop open, community-driven AI safety technologies aimed at addressing AI-related online risks and harms. Experts working at the forefront of policy, research, and product implementation gathered for intensive discussion and development around the use of safety models and related challenges. About 75 people from tech companies, research institutions, and nonprofit organizations, mainly from the United States, participated on the day.
Verification and implementation work was carried out using several open safety models, starting with "gpt-oss-safeguard," the open-weight safety reasoning model released by OpenAI, with projects proceeding across the following three tracks.
- Policy Development: Verifying and improving policy using open safety models
- Model Testing: Practical evaluation, including model performance and cost
- Real-World Applications: Verifying integration into products and workflows assuming real-world operation
Elith took part in this event as the only participating company from Japan, joining hands-on discussion and verification around international knowledge on implementing and evaluating AI safety models, and collaborative efforts leveraging an open technology foundation.
We participated in Track 2 (Model Testing) and Track 3 (Interpretability / Token-level Analysis), conducting technical verification aimed at understanding the behavior of safety models. Targeting gpt-oss-safeguard, we analyzed the factors influencing its judgments and their relationship to policy, and organized the results into a shareable form.
Elith posted the technical outcomes of this hackathon to the official discussions (#39, #40) on the ROOST Model Community, making them public to the international AI safety implementation community.
- Track 2 (Model Testing) — Practical Evaluation of gpt-oss-safeguard:
- Using a multi-layered evaluation pipeline, Elith systematically designed and evaluated 364 adversarial prompts. This quantitatively revealed a tendency for detection failures (bypasses) against gpt-oss-safeguard-20B to occur notably in the Fraud and Malware categories, demonstrating attack patterns expected in real-world operating environments and the model's vulnerability to them.
- Track 3 (Interpretability / Token-level Analysis) — Visualizing the Model's Internal Behavior:
- To understand the internal representations that contribute to the model's safety judgments, we implemented token-level attention-weight analysis using a custom API. This visualizes which tokens are most strongly involved in safety judgments, presenting a technical approach that deepens the interpretability of the "reasons" why particular bypasses occur.
These posts are more than a simple report of results — they represent an ambitious technical achievement, verifying and sharing at an international standard level the implementation risks and behavior of open safety models from the two perspectives of model evaluation and interpretability.
Through this event, Elith shared and absorbed international, hands-on knowledge about the potential and limitations of using safety models in real implementation settings, and the relationship between policy design and model behavior.
We have renewed our recognition that an approach in which diverse organizations collaborate around an open technology foundation to advance AI safety will be essential for the future real-world deployment of AI.
Elith will continue to work with partners in Japan and abroad on initiatives spanning research, implementation, and social responsibility in the fields of generative AI and AI safety.
Event overview
- Name: Open Safeguard Hackathon
- Date: December 8, 2025
- Location: San Francisco, USA
- Organizers: ROOST, Hugging Face, OpenAI
Related information
- Elith's submission (GitHub):
- https://github.com/NaoyaTakashima/attention-safety-guard-apihttps://github.com/roostorg/model-community/discussions/39https://github.com/roostorg/model-community/discussions/40


