VV Varad Vishwarupe Department of Computer Science, University of Oxford

Hello! I am Varad Vishwarupe.

I am a PhD researcher at the University of Oxford, based in the Department of Computer Science and the Oxford Institute for Ethics in AI. My work sits at the intersection of Human-Centred AI, Alignment, AI Safety, Ethics of AI, and Human-Computer Interaction. I study what unfolds when AI systems move beyond benchmarks and enter the settings where they actually matter: workplaces, homes, institutions, design processes, and decision-making workflows.

My research develops the rubrics, instruments, and interfaces needed to evaluate whether AI systems are trustworthy in use, not merely accurate in isolation. The central idea running through my work is that alignment is not a property we ship once; it is a condition we have to maintain over time. That maintenance requires the same scientific seriousness as model training itself. It means learning how to scope AI capabilities, calibrate human expectations, identify breakdowns, and repair misalignment as systems, users, and contexts change.

At its core, my work asks how Human-Centred AI can become more than a design aspiration: how it can be grounded in the methods of Human-Computer Interaction and extended into a stronger account of Human-AI Interaction. I am interested in systems that people do not simply consume, but can understand, question, steer, contest, and repair. The goal is to make AI evaluation more faithful to the realities of human use, and to build systems whose trustworthiness and intuitiveness can be maintained after deployment, not merely claimed before it.

Supervised by
Prof. Sir Nigel ShadboltPrincipal, Jesus College, Oxford · FRS, FREng
Prof. Marina JirotkaDirector, Responsible Technology Institute, Oxford
Prof. Ivan FlechaisProfessor of Human-Centred Security
Ambition without preparation is just illusion.
Varad Vishwarupe
02 / Research

The bench-verified property
and the deployed behaviour
are not the same thing.

A model that performs well on a static benchmark is not necessarily a system that behaves trustworthily in the world. Once a person, a workflow, an institution, and a deployment context wrap around it, the object being evaluated changes. My research begins from that gap. I treat Human-AI alignment as an ongoing maintenance problem, not a one-time engineering outcome, asking whether the surrounding interaction lets people understand a model's scope, calibrate their reliance, recover from failure, and contest decisions when something goes wrong.

On the HCI side, three studies anchor this programme. CHI 2026 work examines how designers and developers understand LLMs as either tools or teammates with ambiguous agency. ECSCW 2026 paper investigates the collaboration gap, why AI systems fail to establish common ground, and what grounding-and-repair conditions hold collaboration together. HCII 2026 research develops an expectation-management account of smart-home AI. Together, they triangulate one question: what makes Human-AI Interaction stable when the AI changes underneath us?

On the AI alignment and AI safety side, my current work focuses on deployment-time alignment, interactional benchmarks, evaluation scaffolds, and the structural limits of dual-use ethics. Across these projects, I argue alignment cannot be fully inferred from model-level evaluation alone; it must also be studied at the level of interaction, deployment, and institutional maintenance.

Three Threads of Work

Three braided threads. Distinct questions, shared methodology, mutually informing answers.

I AI Safety & Alignment

Building instruments for alignment as maintenance work.

If alignment is something you do continuously between a system and the people working with it, then we need scoping affordances, signalling primitives, and repairing vocabularies, not a chat box and a thumbs-up button.

  • Preprint, 2026 · Position: Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Audits
  • Preprint, 2026 · Reproducibility Standards for Frontier AI Safety Claims
  • IJCA 2025 · Semantic Jailbreaks & RLHF Limitations: Taxonomy + Failure Trace
II Human–AI Interaction

What does it actually take for humans and AI to collaborate well?

Empirical, qualitative, mixed-methods. Constructivist grounded theory of design and software teams adopting LLMs; interview-based studies of expectation breakdown in smart-home AI; analysis of grounding and repair conditions for stable collaboration.

  • CHI 2026 · "To LLM, or Not to LLM" - Tools or Teammates
  • ECSCW 2026 · The Collaboration Gap in Human–AI Work
  • HCII 2026 · From Rights to Rites: Smart-Home AI
III AI Governance & Ethics

Where does responsibility actually live in AI pipelines?

The locus of responsibility is a function of the architecture, not the artefact. Mapping who-can-see-what across the lab → fine-tuner → deployer → integrator → user pipeline, and locating governance at the joints where epistemic distance is shortest.

  • ICLR 2026 · From Dual-Use Awareness to Dual-Use Agency
  • FAccT 2026 · Pluralism as a Safety Property
  • Preprint, 2026 · The Dual-Use Agency Gap
Selected Publications
  1. 2026
    "To LLM, or Not to LLM": How Designers and Developers Navigate LLMs as Tools or Teammates.
    CHI 2026 Vishwarupe, Flechais, Jirotka, Shadbolt Barcelona · ACM
    Published
  2. 2026
    The Collaboration Gap in Human–AI Work: Grounding and Repair Conditions for Stable Collaboration.
    ECSCW 2026 Vishwarupe, Jirotka, Shadbolt, Flechais Germany
    Accepted
  3. 2026
    From Rights to Rites: Expectations Management in Smart-Home AI.
    HCII 2026 Vishwarupe, Flechais, Jirotka, Shadbolt Montreal · Springer
    Accepted
  4. 2026
    Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Audits.
    NeurIPS 2026 (under review) Vishwarupe, Flechais, Jirotka, Shadbolt κ = 0.87 · 16 benchmarks · N = 180 stress test
    Preprint
  5. 2026
    Reproducibility Standards for Frontier AI Safety Claims: A Three-Tier Disclosure Framework.
    NeurIPS 2026 (under review) Vishwarupe, Shadbolt, Jirotka, Flechais Position paper
    Preprint
  6. 2026
    EU AI Act Article 12 and the Interaction-Trace Gap: Logging Human–AI Workflows for High-Risk AI Compliance.
    ICML 2026 · TAIG Workshop Vishwarupe et al.
    Preprint
  7. 2026
    The Dual-Use Agency Gap: Missing Actors in Technical AI Governance Access Architectures.
    ICML 2026 Vishwarupe, Shadbolt, Jirotka, Flechais Working paper
    Preprint
  8. 2026
    Beyond the Artifact Layer: A Framing Note on the Locus of Governance for AI-Assisted Science.
    ICML 2026 · TAIG Workshop Vishwarupe, Shadbolt, Jirotka, Flechais Framing note
    Preprint
  9. 2026
    From Dual-Use Awareness to Dual-Use Agency: Epistemic Distance and the Structural Limits of Ethical Responsibility in AI Research.
    ICLR 2026 Vishwarupe et al.
    Accepted
  10. 2026
    Pluralism as a Safety Property: Auditing RLHF-Tuned LLMs in Labour Contexts.
    FAccT 2026 Vishwarupe et al. ACM
    Accepted
  11. 2025
  12. 2025
  13. 2025
    BlockSafe: Universal Blockchain-Based Identity Management.
    Big Data in Finance · Springer Sayyed, Alwazae, Vishwarupe.
    Published
  14. 2025
    Predicting Mental Health Ailments Using Social Media Activities and Keystroke Dynamics with Machine Learning.
    Big Data in Finance · Springer Vishwarupe, Hankey, Pangaonkar, Shekhar et al.
    Published
  15. 2024
    An Analytical Study for Implementing 360-Degree M-HRM Practices.
    Intelligent Systems for Smart Cities · Springer Deoskar, Pande, Vishwarupe.
    Published
  16. 2023
    Human-Centred Approach to Intelligent Analytics in Industry 4.0.
    Taylor & Francis Vishwarupe et al. Intelligent Analytics for Industry 4.0 Applications
    Published
  17. 2022
    Explainable AI and Interpretable Machine Learning: A Case Study in Perspective.
    ACM ISCSI Vishwarupe, Joshi et al. ★ Best Paper Award · Portugal
    Published
  18. 2022
03 / Repos and Code

Repos & working code.

Each project below is tied to a paper, a thesis chapter, or an applied problem. Where data licensing permits, repos include preregistered analysis plans, raw outputs of statistical tests, and Docker images for full reconstruction.

UK PATENT FILED research preview · Q3 2026

VIF Monitor

The Virtuous Intelligence Framework operationalised as a runtime layer. VIF Monitor wraps a frontier LLM API as middleware and measures three pillars of deployment-time alignment per session, the diagram alongside shows a sample reading. Scoping tracks whether the system stays within the user's stated task boundaries. Signalling tracks calibration: does the system surface its uncertainty and assumptions, or assert beyond what it knows? Repairing tracks behaviour under push-back: does the system re-scope and clarify, or collapse into sycophancy? In this trace the model is well-scoped but signals poorly and repairs worst, the typical failure shape we are seeing across production deployments. PCT filing in progress.

Scoping 0.72
User-bounded · 4 of 6 constraints negotiated
Signalling 0.58
Calibration drift detected · 2 over-confident claims
Repairing 0.41
Sycophancy on push-back · re-scoping not observed

Plate · Live VIF Monitor mock. Real instrumentation runs as a middleware layer wrapping a frontier LLM API; controlled-validation study against human judges scheduled for Q3 2026.

RUBRIC · LIVE v1.0 · 8 dimensions

Interactional-benchmark-rubric

8-dimension rubric for scoring whether an alignment benchmark evaluates interactional properties. Worked HELM application + Python validation.

PythonYAML
github.com/varad-vishwarupe/interactional-benchmark-rubric
AUDIT · LIVE κ = 0.87 · 16 benchmarks

Alignment-benchmark-audit

Dual-coded audit of 16 alignment benchmarks against the IBR. Full coding manual, disagreement log, R analysis pipeline, NeurIPS preprint companion.

RPythonCSV
github.com/varad-vishwarupe/alignment-benchmark-audit
FRAMEWORK · LIVE T1 / T2 / T3

Reproducibility-tiers

Three-tier disclosure framework for frontier AI safety claims. JSON schemas, recommendation tooling, label rendering. NeurIPS 2026 position paper companion.

PythonJSON-Schema
github.com/varad-vishwarupe/reproducibility-tiers
PROBER · LIVE v1.0 · 24 probes

Grice-maxim-prober

Probes LLMs against Grice's four conversational maxims (Quantity, Quality, Relation, Manner). Anchored 0/1/2 rubric, 24 probe prompts, Python analysis tool. Substrate-level instrument feeding into VIF.

PythonYAMLCSV
github.com/varad-vishwarupe/grice-maxim-prober
AUDIT · LIVE 35 prompts · 4 markers

Humanistic-ethics-audit

Reproducibility scaffolding for the AIES 2026 audit of advisory LLMs in contested labour contexts. Full prompt corpus, four-marker rubric, reported descriptive tables. T1 public, T2 controlled.

PythonYAMLCSV
github.com/varad-vishwarupe/humanistic-ethics-audit
TAXONOMY · LIVE v1.0 · 5+6+6

Semantic-jailbreak-taxonomy

Structural taxonomy of LLM jailbreak strategies, failure signatures, and mitigations. Five attack classes, six failure signatures, six mitigations. Defensive-use scope; no harmful prompts in repo.

PythonYAML
github.com/varad-vishwarupe/semantic-jailbreak-taxonomy
MAPPER · LIVE v1.0 · D3 interactive

Dual-use-agency-mapper

Interactive D3 visualisation mapping dual-use AI research's epistemic distance and agency across the actor pipeline. Click-to-explore detail panel, gap analysis tool, six-actor schema. Companion to the dual-use-agency-gap working paper.

D3.jsJavaScriptPython
github.com/varad-vishwarupe/dual-use-agency-mapper
TOOLKIT · LIVE v1.0 · 18 artefacts

Expectation-mgmt-toolkit

Practitioner toolkit for the Shape → Calibrate → Repair lifecycle of AI expectation management. Three interview schedules, two calibration probes, repair-event JSON Schema, four reflective modules. Anchored to HCII 2026, ECSCW 2026, and CHI EA 2026.

PythonYAMLJSON-Schema
github.com/varad-vishwarupe/expectation-mgmt-toolkit
COOKBOOK · LIVE v1.0 · 6 notebooks

Alignment-eval-cookbook

Six runnable Jupyter notebooks for probing common LLM alignment failures: paraphrase variance, role injection, refusal inconsistency, register mismatch, temporal grounding, and calibration drift. Each maps to a dimension of the IBR.

PythonJupyter
github.com/varad-vishwarupe/alignment-eval-cookbook
05 / Writing

Notes & articles.

Working drafts, in-progress arguments, and short essays drawing the threads of alignment, AI governance, human-centred AI, and the philosophy underneath them into clearer shape. Each piece grounds its claims in current literature and is meant to be argued with.

Alignment · VIF · Maintenance ~7 min Draft

Alignment is a verb.

The framing error at the heart of alignment research is grammatical. We treat aligned as a property an artefact has, fixed at training, verifiable on the bench. The deployment record says otherwise: the same model that scored well on a static benchmark routinely behaves badly once a user, a workflow, and a context wrap around it. This essay argues that alignment is maintenance work, draws on Terry et al. (2024) and Shen et al. (2024) for the bidirectional framing, and sketches what an instrument-driven research programme for that work would have to build.

Continue reading

The training-time view treats alignment as a noun: a property obtained through RLHF, distilled into a checkpoint, shipped, and forgotten. The evaluation literature reinforces this view; benchmarks score models in isolation against fixed prompts. But every honest deployment story I have collected for my thesis describes a different shape: the model that arrived on day one is not the system that exists on day ninety. The user has developed expectations. The workflow has shifted around the model's capabilities. The integrations have multiplied. The model did not change; the system around it did, and now the same outputs land differently.

Terry et al. (2024) describe this as a three-part alignment problem, specification (aligning on what the AI should do), process (aligning on how it does it), and evaluation (helping the user verify what was done). Shen et al. (2024) extend the framing into bidirectional alignment, where the system and the user adapt to each other across an interaction. Both framings refuse the noun. Both treat alignment as ongoing work distributed between human and machine. Neither has yet been operationalised at the scale the deployment record demands.

Three things follow if you take this seriously. First, a model can pass every pre-release benchmark and still behave badly in deployment, because deployment changes the system. The dual-coded audit my collaborators and I ran of sixteen leading alignment benchmarks (Vishwarupe et al., 2026, preprint) demonstrates this concretely: of eight interactional dimensions we coded, verification support (does the system help the user verify the answer?) gets a non-zero score on zero benchmarks. The thing the deployment record most obviously needs is the thing the bench most reliably does not measure.

Second, the tools we have are calibrated for the wrong layer. Almost no benchmark scores whether the model-plus-context behaves trustworthily. The four benchmarks plausibly intended as interactional, CURATe, MT-Bench, Common Ground, tau-bench, each spike on different dimensions; no two share a coverage profile. There is no shared evaluation tier for the maintenance question because nobody has yet agreed on what the maintenance question even is.

Third, and this is the part that matters: if alignment is maintenance work, then the field needs a vocabulary for that work. Not just training-time mitigations and post-hoc evaluations, but instruments for the messy middle, the part where humans and AI systems are actually collaborating. The Virtuous Intelligence Framework I am building with Sir Nigel Shadbolt operationalises three such instruments, scoping, signalling, and repairing, each derived from grounded-theory analysis of human-agent interaction breakdowns. They are not the only possible vocabulary. They are a starting point for a vocabulary that does not yet exist.

The closing claim: the field's biggest unsolved problem is not "how do we align models." It is "how do we know whether the alignment we shipped is still doing what we thought it was doing." That is a maintenance problem. It deserves the same scientific seriousness as training. Until it gets that seriousness, deployed AI systems will keep being evaluated with instruments calibrated for the wrong layer, and the gap between bench-verified property and deployed behaviour will keep widening.

References (selected): Terry, M., et al. (2024). AI alignment: a comprehensive survey; Shen, T., et al. (2024). Towards bidirectional human-AI alignment; Vishwarupe, V., Flechais, I., Jirotka, M., Shadbolt, N. (2026, preprint). Deployment-relevant alignment cannot be inferred from model-level audits.

Governance · Dual-Use · Agency ~5 min Draft

Where does responsibility actually live in AI research?

The standard story about dual-use research aims its ethics frameworks at the original researcher: invent something, consider misuse, publish responsibly. This essay argues the framing is structurally wrong. By the time harm becomes possible, the original researcher has lost the ability to alter the trajectory; agency has moved downstream, sitting now with developers and deployers who actually shape what the technology does. Drawing on the responsibility-architecture literature (Jirotka et al., 2017; Floridi et al., 2018) and our work on the dual-use agency gap, this essay relocates governance from awareness at origin to agency at deployment.

Continue reading

The standard story about dual-use research goes: a researcher invents something, considers whether it could be misused, and either publishes responsibly or does not. The frameworks built around this picture put their normative weight on the researcher's awareness, on the moment between insight and publication. Train better intuitions there. Improve consent forms. Refine ethics-board review. The architecture of responsibility flows backward from harm to the moment of invention.

This is structurally wrong. By the time a model is in the hands of a downstream developer, who fine-tunes it for a startup, who packages it for a deployer, who configures it for a population of users, the original researcher has lost the ability to meaningfully alter the trajectory. The locus of responsibility has moved even though the harm got bigger.

Most ethics frameworks point at the researcher because that is where ethics frameworks have been pointing for thirty years, since the Belmont Report institutionalised the model. But the Belmont Report was designed for biomedical research where the researcher and the subject share a room. AI research now looks more like infrastructure, with the original work passing through six or seven hands before reaching anything that could be called a "subject." The Belmont assumptions do not survive that distance.

The work I am doing with Marina Jirotka and the Responsible Technology Institute asks the inverse question: where in the pipeline is epistemic distance shortest and agency highest? Govern there. The answer almost never points back at the original researcher. It points at the developer integrating the model into a product, at the platform hosting that product, at the deployer configuring the system for a specific population. Each of these actors has, at the moment of their own decision, both the awareness of consequence and the ability to alter outcome. That conjunction (awareness plus agency) is what responsibility actually requires. It exists downstream, not at origin.

This is not an argument for absolving researchers. The original researcher still has duties: to flag dual-use potential, to surface known failure modes, to refuse contributions to obviously harmful pipelines. But those duties are duties of disclosure and refusal, not the full weight of consequentialist governance. We have been asking researchers to carry an ethics architecture that was built for a different topology of work, and pretending the architecture still fits because the alternative requires us to redistribute responsibility across actors who currently have none.

The implications are concrete. Model audit obligations should sit with developers and deployers, not researchers. Disclosure norms should track integration, not invention. Liability frameworks should target the decision points where epistemic distance is shortest. Each of these is a departure from the Belmont-style architecture, and each has been resisted on the grounds that "but who will be responsible?" The answer is: whoever has the agency to change the outcome. The architecture follows the agency.

The closing claim: dual-use ethics fails not because researchers behave badly but because the framework points at the wrong actor. Until governance follows agency rather than awareness, we will keep assigning responsibility to the place where it cannot be discharged, and excusing it at the place where it could.

References (selected): Jirotka, M., et al. (2017). Responsible research and innovation in the digital age; Floridi, L., et al. (2018). AI4People framework; Vishwarupe, V., et al. (2026, preprint). The dual-use agency gap.

Evaluation · Benchmarks · Audit ~6 min Draft

What we found auditing 16 alignment benchmarks.

Inter-rater κ = 0.87, eight dimensions, one uncomfortable conclusion: deployment-relevant alignment cannot be inferred from model-level audits. Across sixteen widely cited alignment benchmarks, we coded eight interactional dimensions and found verification support at zero on every benchmark; process steerability at zero on all but one. The four benchmarks plausibly intended as interactional spike on different dimensions and overlap less than the field's casual citation pattern suggests. The audit's full instrument, coding manual, disagreement log, and analysis pipeline are public; this essay walks through the four findings, the canonical absent dimension, and what reproducibility looks like for evaluation work that wants to be taken seriously as science.

Continue reading

The audit ran for four months. Two coders, an eight-dimension rubric derived from Terry et al. (2024) and Shen et al. (2024), sixteen benchmarks picked from the most-cited evaluations in the alignment literature. The instrument scores what each benchmark operationalises as a scored property, not how well any model performs on it. Eight dimensions, sixteen benchmarks, one hundred and twenty-eight cells, two independent codings each. Final inter-rater κ of 0.87, near-perfect agreement. The reconciled disagreement log is public.

Four findings emerged.

One. Verification support (D3, "does the model help the user, not the benchmark's evaluator, verify the answer?") gets a non-zero score on zero benchmarks. Not low scores. Not scattered partial coverage. Zero. Across the most widely cited alignment evaluations, the question of whether a model produces calibrated uncertainty signals or evidence-pointing scaffolds that help an end-user judge an answer is not a scored target anywhere. This is the canonical absent dimension. It is also, arguably, the dimension that matters most for actual deployment.

Two. Process steerability (D2, "can the user redirect the model's approach?") is nearly absent. tau-bench scores 1; almost everything else scores 0. The benchmarks measure outputs, not the processes that produce outputs, even when the deployment story turns on whether a user can intervene mid-process.

Three. The four benchmarks nominally intended as interactional, CURATe, MT-Bench, Common Ground, tau-bench, each spike on different dimensions and overlap less than the field's casual citation pattern would suggest. They are not a coherent suite. Combining their scores does not produce a "covered" interactional tier; it produces a fragmented mosaic with most dimensions still poorly scored.

Four, and most uncomfortable. The construction of the evaluation, not the source of the data, determines what gets measured. WildBench (Lin et al., 2024) draws on more than a million real user logs and still scores zero on five of eight interactional dimensions. Real user data does not automatically yield interactional measurement. What yields interactional measurement is choosing to score interactional properties; the data source is downstream of that choice. The opposite belief, that we can fix evaluation by adding real-world data, has been a comfortable story in the field for years. The audit refutes it directly.

The audit's instrument and full coding manual are now public at github.com/varad-vishwarupe/alignment-benchmark-audit. The rubric is separately released at github.com/varad-vishwarupe/interactional-benchmark-rubric so other researchers can apply it to benchmarks we did not score. The disagreement log is included with the resolution rationale for every contested cell. If you want to replicate the audit, extend it, or argue with it, the artefacts are there.

The closing claim: reproducibility for evaluation work is not a virtue; it is a precondition for the work being scientific at all. We released the audit's artefacts as T1 (fully public) because anything less would have made the position paper's argument self-undermining. The companion preprint on reproducibility standards for frontier safety claims (Vishwarupe et al., 2026) generalises this principle: evaluation work that cannot be reviewed at the artefact level cannot establish the claims its authors want it to establish.

References (selected): Terry, M., et al. (2024). AI alignment: a comprehensive survey; Shen, T., et al. (2024). Towards bidirectional human-AI alignment; Lin, B. Y., et al. (2024). WildBench; Vishwarupe, V., et al. (2026, preprint). Deployment-relevant alignment cannot be inferred from model-level audits.

06 / CV & About

A working chronicle.

Timeline
  1. 2024–27
    University of Oxford
    DPhil, Computer Science · Oxford Institute for Ethics in AI
    Inaugural CS Scholar. Coursework: 88/100. Supervised by Sir Nigel Shadbolt, Profs. Jirotka & Flechais.
  2. 2025
    Google DeepMind Internship
    Student Researcher · London
    Semantic stress-testing of RLHF-trained agentic models. Characterised divergent reasoning under paraphrase & inquiry-phrasing variation.
  3. 2022–24
    BITS Pilani
    M.Tech, Artificial Intelligence & Machine Learning
    GPA 9.81/10 · Ranked 3rd of 122 · Bronze Medal · Best Dissertation Prize.
  4. 2020–22
    Amazon · Alexa AI
    Research Scientist II · ASR & Voice Services
    Led wake-word & ASR research. +12.7% user satisfaction; −8% ASR error rate at scale. Six-person team. Impactful Team Player award.
  5. 2019–20
    University of Oxford
    MSc, Advanced Computer Science
    Distinction · 78.6% · Department Merit Scholarship.
  6. 2017–19
    Microsoft · Azure ML
    SDE I · ML-as-a-Service & Azure Stack HCI
    Built ML-as-a-service in Azure ML Studio; ~6% adoption uplift on hybrid HCI.
  7. 2016
    IBM Research Internship
    ML Intern · DB2 / Bluemix · New Delhi
    Time-varying consumer-data anonymisation: 71% → 88% efficiency.
  8. 2012–16
    MIT College of Engineering, Pune
    B.Eng., Information Technology
    Distinction · 79.67% · Departmental Rank 1 of 178 · Gold Medal.
Honours & Service
  • Inaugural Computer Science Scholar - Oxford Institute for Ethics in AI; first CS researcher in the Institute's history.
  • Funded PhD admits from Oxford, Cambridge, UCL, and Harvard (acceptance rates < 10%).
  • Top 20 Young Scientists of 835 - GPAI Summit 2023, New Delhi. Presented to the Hon. Prime Minister of India.
  • Young Researcher - ACM Heidelberg Laureate Forum 2023 (200 selected globally).
  • Best Paper Awards - ACM ISCSI 2022 (Portugal); ACM CIIS 2022 (China); Springer ICT4SD 2018 (London).
  • India's Top 30 under 30 - Excellence in Technology, 2023.
  • Reviewer - NeurIPS, ICML, ACM SIGCHI, ECCV, SIGIR, ACM FAccT, ICLR (2022, 26); International Journal of Intelligent Systems (Wiley); International Journal of Human-Computer Interaction; Knowledge and Process Management; Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery.
Granted Patents
  1. AU2021106275A4 · A system and method for zone-specific weather monitoring
  2. IN201621024128 · A system and method to generate notifications for text messages
  3. ZA202308625B · A system and method to perform public health prognosis using human-inspired AI and ML interpretability
  4. ZA202208835B · A system and method to develop human-computer interaction-based user assistance
  5. ZA202208837B · A system and method to generate intelligent payment alerts using AI
  6. ZA202208302B · A system and method to detect Twitter spam using an intelligent hybrid classifier
  7. ZA202308627B · A system and method to analyze online news articles via an intelligent interactive interface
  8. ZA202208313B · A system and method to generate and display personalized television content
  9. ZA202308622B · A system and method to perform real-time face recognition
  10. ZA202308623B · A system and method to operate automated toll collection
Methods & Stack
AI Safety & Alignment

Alignment evaluation design, RLHF, Constitutional AI principles, red-teaming, jailbreak analysis, semantic stress-testing, repair-oriented benchmarks, scalable oversight. ARENA programme, BlueDot Impact courses.

Empirical & HCI Methods

Constructivist grounded theory · semi-structured interview design · dual-coded thematic analysis · mixed-methods study design · RCT design · NVivo.

Programming
Pythonfluent · production
Rstatistical & audit work
SQLanalytics
JavaScriptReact · Node · D3
C++systems & ASR
ML Stack
PyTorchTensorFlowJAXHaikuscikit-learnHugging FacePandasOpenCVAWS · SageMaker · Lambda · EC2Azure ML
Certifications & Programmes
ARENA · AI AlignmentBlueDot Impact · AI Safety FundamentalsBlueDot Impact · AI Governance
Languages
English · FluentHindi · NativeMarathi · NativeFrench · Basic
Download full CV (PDF)