MIRAGE benchmark accuracy
Built an agentic oncology intelligence system with an evidence-weighted knowledge graph and auditable retrieval, 7.6 points above baseline.
- ~18KGraph entities
- ~95KRelationships
Applied Scientist · NLP, Retrieval & LLM EvaluationNew York, NY
Applied scientist and ML engineer building retrieval, classification, and LLM evaluation systems. I pair production ownership with rigorous experimentation, model validation, and research on robustness and efficiency.

Measured results across production ML, retrieval, knowledge systems, and model safety.
Built an agentic oncology intelligence system with an evidence-weighted knowledge graph and auditable retrieval, 7.6 points above baseline.
Built a schema-validated document-intelligence pipeline for scanned PDFs, forms, and tables across 800 documents.
Developed a training-time safety defense for an open-weight LLM while preserving baseline behavior.
Built the evaluation backend for a production measurement platform spanning six LLMs and more than 10,000 prompts each week.
Production ownership, careful evaluation, and research depth across the ML lifecycle.
I build hybrid retrieval, ranking, and multi-agent systems with traceable evidence, calibrated routing, and latency and cost measured against real constraints.
I design benchmarks, confidence thresholds, and out-of-domain checks that expose failure modes and make model quality useful for operational decisions.
I study how model behavior is represented, attacked, and improved. Current work spans refusal robustness, activation-level interpretability, and tokenizer optimization.
Production ML work that improved quality, latency, and the decisions teams could make.
Three focused examples of applied systems and research, each framed by the problem, method, and result.
Developed a training-time defense and evaluation harness for an open-weight LLM, then stress-tested whether refusal behavior could be removed through low-rank linear ablations.
Increased the rank needed to break refusal from 1 to at least 16 while preserving baseline model behavior.
Developed a Gurobi mixed-integer optimization approach for tokenizer vocabulary selection that matches the greedy decoding rule used by WordPiece.
Reduced token count by up to 0.78% versus BPE at the same vocabulary size and closed 89.6-99.4% of the remaining compression gap. Accepted at COLM 2026 Tokshop.
Built a fully local physical-therapy coach that combines Qwen2.5-VL form feedback, MediaPipe repetition tracking, range-of-motion analysis, and spoken guidance in the browser.
Top-8 of 30 teams at the Dell × NVIDIA Hackathon 2026 (NYU CDS).
Consistently suppressed targeted activations, including fully turning off one causally validated feature, but found that readable activations do not always control behavior.
Improved ROC-AUC by 6.2 percentage points over CNN baselines and shipped an interactive 3D demo.
Evaluated motion forecasts across 1.6, 3.2, and 5.0 second horizons using deployment-focused trajectory metrics instead of single-step error alone.
Ran customer discovery and MVP scoping to validate pain points across clinical research operations and architecture / building-code review.
Selected publications, education, and the technical foundation behind the work.
Joint Optimization for Greedy Longest-match Tokenization
Developed a mixed-integer optimization approach that matches the greedy decoding rule used at deployment, reducing token count by up to 0.78% versus BPE at the same vocabulary size.
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Showed that fluent prompts can suppress five types of internal model features without changing model weights, while control experiments revealed that readable activations do not always control behavior.
Validity of Machine Learning-Based COVID-19 Prediction
Validated seven classifiers on 195,000 clinical records, measured approximately 20% AUROC degradation under cross-continental distribution shift, and released an open-source evaluation toolkit.
Workshop paper · AAAI Deployable AI · 2023
Review article
22nd Int'l Conference on Bioinformatics, Brisbane · Nov 2023
I am open to applied scientist, machine learning engineer, data scientist, and research engineer roles.
dm6262@nyu.edu