Hiring We're hiring for research scientists and research engineers, in Singapore and remotely! See all roles →

Frontier AI safety,
built in Asia.

We're an independent research and evaluations lab working to ensure that frontier AI can be safely deployed at scale.

What we do

Evaluations

Independent third-party evaluations and model safety reports for frontier AI systems.

Our methodology is aligned with the systemic-risk evaluation obligations of the EU AI Act Code of Practice, with coverage across CBRN, cyber, harmful manipulation, and loss-of-control risks.

Research

Empirical research on the failure modes that emerge as frontier AI systems become more capable.

Our research develops the science of safety and new methodologies for evaluating misalignment and loss-of-control risks, with focus on the failures that emerge in agentic systems.

Collaborations

We work closely with academic labs, frontier safety organisations, and standards bodies.

Our collaborations span joint evaluations, shared methodology, and the technical work needed to keep frontier AI safety robust across institutions and borders. Reach out if you’d like to collaborate on a project.

2026 Roadmap

Research timeline

Updated 02 May 2026
Q2 April – June
EU AI Act compliant model safety cards

Independent safety evaluation reports on frontier models, combining established benchmarks with follow-up experiments that probe specific risks in greater depth.

Intended as a reference template for independent safety reporting under emerging regulatory regimes.

In progress
Q3 July – September
Testing and evaluating safety mitigations

Independent stress tests of frontier safety mitigations such as chain-of-thought monitoring and constitutional classifiers.

We plan to publish the methodology and findings as practical guidebooks, aimed to lower the cost of adoption for labs that want to deploy these mitigations without rebuilding the evaluation work themselves.

Planned
Q4 October – December
Misalignment and loss-of-control research

Empirical research on misalignment and loss-of-control risks in frontier models. Threads include evaluation awareness, sandbagging, and shutdown avoidance.

Planned