Artificial intelligence is rapidly becoming part of everyday life, powering everything from search engines and healthcare tools to financial services and online education. As these systems become more capable, they are also becoming more influential, making decisions that can affect people's opportunities, privacy, and security. Ensuring that AI systems are accurate, secure, transparent, and worthy of public trust has become one of the defining challenges of modern computing.
At CyLab, our researchers are tackling this challenge from multiple directions. Rather than focusing only on making AI systems more powerful, CyLab researchers are working to make them more trustworthy, identifying hidden vulnerabilities, measuring real-world risks, improving fairness, and creating practical methods for evaluating AI before and after it is deployed. Together, their work is helping shape a future in which AI systems are not only more capable, but also more dependable and accountable.
Finding AI's Weaknesses Before Attackers Do
One of the most important steps toward building trustworthy AI is understanding how these systems fail. CyLab researchers Matt Fredrikson and Zico Kolter have become international leaders in identifying the vulnerabilities of large language models (LLMs) and other machine learning systems, demonstrating that even today's most advanced AI models can often be manipulated in unexpected ways.
Their research has shown that AI models can be fooled into producing inaccurate, harmful, or confidential information through carefully crafted inputs known as adversarial attacks. By exposing these weaknesses under controlled research settings, Fredrikson and Kolter help developers better understand how AI systems behave when confronted with malicious users. Their research led them to publicly launch Gray Swan in 2024, an enterprise security platform that provides automated adversarial testing, continuous red teaming, and runtime AI security for enterprises deploying AI and the frontier labs building it.
In August of 2024, Kolter joined the board of OpenAI and became the chair of its Safety and Security Committee, overseeing model development governance and release safety. The pair were also contributing authors to the definitive paper on Indirect Prompt Injections, “How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition,” in March of 2026
Rather than undermining confidence in AI, Fredrikson and Kolter's work strengthens it by revealing vulnerabilities before they can be exploited in the real world and informing the development of more robust defenses. Their broader research has helped establish the field of adversarial machine learning, laying the scientific foundation for building AI systems that remain reliable even in challenging or hostile environments.
Project links:
- Gray Swan
- Podcast: AI Security after Codex and Claude Code
- News story: Researchers discover new vulnerability in large language models
In March of 2026, Hoda Heidari was appointed as one of 40 experts worldwide to the UN’s first Independent International Scientific Panel on AI.
Understanding AI Risks Beyond the Technology
Building trustworthy AI requires more than solving technical problems. Many of the most significant risks emerge from how AI is developed, deployed, and used by people.
CyLab researchers are increasingly examining these broader societal impacts to better understand where AI systems succeed and where they can cause harm.
Hoda Heidari’s research explores how AI systems can be designed to better reflect human values such as fairness, accountability, and transparency. Rather than treating fairness as a purely mathematical problem, her work considers the complex tradeoffs involved when AI systems affect different groups of people.
In their 2025 paper “A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents,” Heidari and collaborators Megan Li, Wendy Bickersteth, Ningjing Tang, Jason Hong, Lorrie Cranor, and Hong Shen analyzed nearly 500 publicly reported incidents involving generative AI to develop one of the most comprehensive taxonomies of real-world AI failures to date. Their findings showed that many harms arise not simply from technical errors, but from the ways AI systems are deployed and used in practice. The research argues that improving AI safety will require stronger governance, greater public understanding of AI, and responsible deployment practices alongside technical advances.
In March of 2026, Heidari was appointed as one of 40 experts worldwide to the UN’s first Independent International Scientific Panel on AI, a body specifically created to translate scientific evidence into information governments can use when developing AI policy. In July of the same year, Heidari and the other panel members released their first Preliminary Report on AI, which synthesizes evidence on issues closely connected to her research, including algorithmic fairness, unequal distribution of AI’s benefits and harms, human rights, accountability, and the challenges governments face in evaluating AI systems. The UN presented the report to governments at its inaugural Global Dialogue on AI Governance in Geneva on July 6–7, 2026, explicitly positioning it as a common scientific foundation for policymaking.
Sauvik Das’s research complements this work by examining how AI creates new privacy and security risks for individuals. His research investigates how increasingly capable AI systems can be used to infer sensitive information, amplify surveillance, or expose users to new forms of manipulation. By identifying these emerging risks before they become widespread, Das helps inform the design of AI technologies that better protect users' privacy while enabling organizations and policymakers to anticipate the unintended consequences of rapidly evolving AI capabilities.
Das has co-authored papers in areas such as AI privacy for practitioners, human-centered adversarial machine learning, and physically-intuitive privacy and security, including Best Paper at the 2024 ACM CHI conference on Human Factors in Computing Systems, “Deepfakes, Phrenology, Surveillance, and More! A Taxonomy of AI Privacy Risks.”
Das’s research on targeted advertising and privacy has helped inform federal debates over government access to commercially collected data. His work was cited in a 2026 ACM policy response to U.S. immigration and homeland security officials, supporting calls for stronger safeguards against repurposing advertising data for government surveillance. That same year, Das was appointed as privacy co-chair of ACM’s U.S. Technology Policy Committee.
Project links:
- UN Independent International Scientific Panel on AI Preliminary Report
- Association for Computing Machinery Response to ICE and Homeland Security Investigations Request for Information from Big Data and Ad Tech Providers
- SPUD (Security, Privacy, Usability and Design) Lab
Putting People at the Center of AI Accountability
Trustworthy AI depends on understanding how people experience these technologies in everyday life. Jason Hong and his collaborators approach AI accountability from a human-centered perspective, recognizing that the people most affected by AI systems often have valuable insights into how those systems succeed or fail.
One example is WeAudit, a framework that enables everyday users to participate in auditing generative AI systems. Instead of relying exclusively on technical experts, WeAudit helps users systematically document problematic AI behaviors and communicate those findings in ways developers can act upon. The project demonstrates that meaningful AI evaluation can come from the lived experiences of diverse users, helping practitioners identify issues that conventional technical testing may overlook.
Hong's team has also explored how to broaden participation in AI accountability. In “Investigating Youth AI Auditing,” researchers showed that teenagers can meaningfully identify bias and other problematic behaviors in AI systems, often recognizing issues adults might miss because of their unique experiences and perspectives.
In another study examining AI in homeless services, the team worked directly with frontline service providers and people experiencing homelessness to understand how AI-assisted decision-making could affect vulnerable populations. Rather than assuming technology alone can solve complex social problems, the research emphasizes designing AI systems with the participation of the communities they are intended to serve.
Hong's work on AI auditing helped inform the National Telecommunications and Information Administration’s 2024 AI Accountability Policy Report, which called for impacted communities to play a role in exposing AI risks and recommended stronger federal support for independent audits and evaluations of high-risk AI systems
Together, these projects illustrate a central principle of CyLab's work: trustworthy AI requires not only better algorithms, but also meaningful engagement with the people whose lives those algorithms affect.
Project links:
- Paper: WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI
- Video: User-driven Auditing & WeAudit - Jason Hong
- Paper: Investigating Youth AI Auditing
- Paper: Understanding Frontline Workers' and Unhoused Individuals' Perspectives on AI Used in Homeless Services
- NTIA 2024 AI Accountability Policy Report