Building
Projects
Open-source tools I build around AI safety, evaluation, and governance. Each links to its repository.
All source code, software, datasets, documentation, and other project materials are independently developed or derived from publicly available resources and are not affiliated with, sponsored by, or endorsed by any individuals or organizations.
- AI Governance & Assurance Library ↗A versioned library of governance and assurance artifacts — frameworks, assessments, testing procedures, checklists, and evidence templates — organized by purpose and mapped across NIST AI RMF, the EU AI Act, OWASP, and SR 26-2, with an overlay for agentic AI.
- GenAI Alignment ↗A scenario library and testing framework for GenAI alignment — objective drift, robustness, and agentic/enterprise risk — tested against Azure OpenAI and Claude, with audit-ready findings mapped to governance frameworks.
- Regulus ↗AI governance standards lookup powered by RAG and knowledge graphs — retrieves applicable risks and cross-referenced guidance across NIST AI RMF, SR 26-2, the EU AI Act, and more.
- LLM Red Teaming ↗A modular toolkit for red teaming LLMs — adversarial attacks, jailbreak evaluation (JailbreakBench), and prompt injection, with pluggable targets and automated judges. Built for AI safety practitioners.
- GenAI Capability Bench ↗A modular benchmark suite evaluating GenAI across accuracy, truthfulness, instruction following, reasoning, RAG, tool use, and agentic task completion.
- Multi-Agent OTel Eval ↗Enterprise-grade GenAI observability and evaluation using OpenTelemetry conventions, with multi-agent orchestration tested on the Mind2Web benchmark.
- RAG Eval Framework ↗Provider-agnostic RAG evaluation benchmarked on HotpotQA — 13 metrics, multi-prompt comparison, failure diagnosis, and an auto-generated audit report.
- Geometric Knowledge Network ↗A lightweight geometric knowledge network that augments vector RAG with document-grounded graph structure, hybrid retrieval, and evaluation.
- ML Validation Framework ↗Interactive ML model validation — explainability, weak spots, robustness, and fairness — widget-driven, no-code demo.