Building
Projects
Open-source tools I build around AI safety, evaluation, and governance. Each links to its repository.
- GenAI Alignment ↗A scenario library and testing framework for GenAI alignment — objective drift, robustness, and agentic/enterprise risk — tested against Azure OpenAI and Claude, with audit-ready findings mapped to governance frameworks.
- Regulus ↗AI governance standards lookup powered by RAG and knowledge graphs — retrieves applicable risks and cross-referenced guidance across NIST AI RMF, SR 26-2, the EU AI Act, and more.
- LLM Red Teaming ↗A modular toolkit for red teaming LLMs — adversarial attacks, jailbreak evaluation (JailbreakBench), and prompt injection, with pluggable targets and automated judges. Built for AI safety practitioners.
- GenAI Capability Bench ↗A modular benchmark suite evaluating GenAI across accuracy, truthfulness, instruction following, reasoning, RAG, tool use, and agentic task completion.
- Multi-Agent OTel Eval ↗Enterprise-grade GenAI observability and evaluation using OpenTelemetry conventions, with multi-agent orchestration tested on the Mind2Web benchmark.
- RAG Eval Framework ↗Provider-agnostic RAG evaluation benchmarked on HotpotQA — 13 metrics, multi-prompt comparison, failure diagnosis, and an auto-generated audit report.
- Geometric Knowledge Network ↗A lightweight geometric knowledge network that augments vector RAG with document-grounded graph structure, hybrid retrieval, and evaluation.
- ML Validation Framework ↗Interactive ML model validation — explainability, weak spots, robustness, and fairness — widget-driven, no-code demo.