Research vision. My research asks how society can delegate consequential decisions and actions to complex software without surrendering the ability to verify, constrain, and correct it. I began with automated accounts on the social Web, moved to distributed and decentralised systems, and now study artificial intelligence systems that generate content, invoke tools, and coordinate with other agents. The technical objects have changed, but the security question has remained stable: when a system mediates human knowledge and agency, what evidence should make us trust it, and what mechanisms let us intervene when that trust is misplaced? I address this question through AI safety, security, and privacy (SSP) under an operational view of AI governance. For me, governance is not a checklist added after a model is built. It is a systems research agenda that connects measurement, certification, policy, runtime enforcement, observability, incident response, and recertification. My goal is to establish an AI assurance research programme that develops these foundations and tests them in partnership with researchers, industry, public institutions, and the people who operate the systems.
My doctoral research began with the gap between the apparent trustworthiness of the social Web and the weak security assumptions beneath it. I built and carefully operated a network of 100 socialbots to test whether malicious automation could reproduce the behaviours that made a social account appear human. Over eight weeks, the bots achieved infiltration rates of up to 80 percent, collected 250 GB of data associated with more than one million users, and faced detection rates below 20 percent. The work received the ACSAC Outstanding Paper Award [1,2] and established an enduring principle for my research: security mechanisms should be evaluated against adversaries that can imitate normal behaviour, not only against convenient anomalies.
I then turned from attack measurement to defence. Existing fake-account detectors assumed that malicious accounts would remain socially isolated or visibly spammy. Our graph-based analysis, which received the ASONAM Best Paper Award [3], showed that these assumptions failed once socialbots formed credible ties [4]. I led the development of Integro, which instead predicted likely victims and incorporated those predictions into an infiltration-resilient ranking scheme. A production evaluation at Tuenti processed a graph of 160 million nodes on 33 commodity machines in under 30 minutes and detected nearly ten times more fake accounts than the deployed baseline. The research appeared at NDSS [5] and in Computers & Security [6], contributed to a U.S. patent application, and demonstrated the value of joining adversarial modelling, distributed systems, and deployment evidence.
These projects shaped three habits that remain central to my work: model the capable adversary; make the security boundary explicit; and validate ideas in the environment where decisions carry costs.
In 2016, I shifted from protecting centralised platforms to studying systems that distribute trust. Socialbots exemplify the Sybil attack: an adversary enters a system under multiple identities and uses them to manipulate its operation. Blockchains address the same underlying challenge from a different architectural direction, combining Sybil-resistant consensus with a durable shared record that does not depend on a single controlling party. To investigate this design space, I founded the Cybersecurity Initiative for Blockchain Research (CIBR) at QCRI and grew it to eight affiliated researchers and engineers.
The programme joined on-chain transaction structure with off-chain evidence. BlockTag, published at IFIP SEC [7], introduced an extensible method for attaching auxiliary evidence to blockchain addresses. Dizzy, published at ARES [8], scaled the collection and analysis of onion services and linked them to blockchain transactions [9,10]. With Toshi and Kansa, these ideas became systems used by the Financial and Electronic Crime Combating Unit at Qatar's Ministry of Interior to investigate cryptocurrency fraud and ransomware. The programme contributed to multi-year research projects, patents, international collaborations, and public-interest deployment.
This phase taught me that provenance is useful only when it supports a decision. A ledger can preserve evidence, but analysts still need defensible methods for linking records, assigning confidence, protecting sensitive data, and explaining conclusions. That insight became the bridge to my current research on generative AI.
I led the attribution subproject for Fanar [11], QCRI's Arabic-centric multimodal generative AI platform. Fanar's opt-in post-generation service decomposes an answer into factual claims, retrieves supporting sources, and revises the response so that claims can be checked against cited evidence. It treats attribution as a system component rather than a stylistic request to the language model. In a related direction, TokenX connects attribution with blockchain-based accounting so that publishers can receive compensation when their openly available content supports a model response. This work brought my earlier interests in open information systems, evidence, and decentralised incentives directly into AI.
Deployment also changed the research question. A model with attribution can still be poisoned, a benchmark can still give a misleading assurance signal, and an agent can still route untrusted content to a privileged tool. Fanar therefore became both a contribution and a catalyst: building an AI platform made clear that we also needed methods to secure it, audit it, and govern changes across its lifecycle.
My current programme studies assurance at four connected layers.
4.1. Models and the AI supply chain
In Poison with Style, accepted at ICML 2026 [12], we showed that a developer's coding style can act as a covert trigger in a poisoned code language model. The attack caused the model to emit code with a CWE-20 vulnerability in 95 percent of triggered cases while reducing pass@1 by less than five percentage points on HumanEval and MBPP. The result challenges evaluations that test utility and backdoors separately: an apparently capable model can carry a latent vulnerability activated by ordinary developer behaviour. I am extending this direction toward AI bills of materials, provenance for models and context, and tests that connect a supply-chain claim to an observable security property.
4.2. Evaluation as a measurement system
I co-lead aiXamine [13], a black-box platform that evaluates trade-offs across LLM safety, security, and privacy. The platform contains more than 45 benchmarks across nine evaluation domains and has run over 5,000 examinations of more than 130 models. The peer-reviewed framework, appearing at RAID 2026, makes cross-dimensional trade-offs visible instead of compressing them into a single score.
Our subsequent reliability audit with Sayf-Eval [14], currently under review, treats a benchmark as a pipeline from dataset and prompt through inference, extraction, scoring, and aggregation. Across eight cybersecurity benchmarks, 48,662 questions, 23 tasks, and ten models, we identified 15 systematic failure modes. One pipeline choice changed a score by more than 80 percentage points; after standardisation, nine of ten models moved at least three ranks on at least one benchmark. This motivates executable reference evaluators and evaluation pipeline cards that record exactly what a score measures.
Static benchmarks also saturate and may enter training data. In SSP-Bench [15], currently under review, we generate fresh, source-grounded items and select them for validity, difficulty, discrimination, novelty, and diversity. Across 24 models and four SSP evaluation domains, the method exposed construct mixing in safety leaderboards and model-family regressions hidden by static scores. This work moves evaluation from repeated benchmark-taking toward measurement science.
4.3. Agentic systems and pre-deployment policy inference
Agentic applications assemble their attack surface from models, tools, prompts, remote services, and delegations. The resulting interaction graph is usually implicit in source code and incomplete at review time. In AgentSA, currently under review, we developed a framework-independent model of these interactions and a static analyser that reconstructs the graph without importing or executing the application. Across seven real systems, six synthetic benchmarks, and 209 manually audited interactions, the analyser achieved 0.99 precision for recoverable edges and produced 187 authorisation rules plus 57 default-deny rules without manual policy authoring.
The next step is to move from recovered topology to least-privilege design: infer candidate access-control policies, identify paths from untrusted inputs to privileged sinks, distinguish statically enforceable edges from runtime-mediated choices, and monitor policy drift as applications evolve across frameworks such as MCP and A2A.
4.4. People who bear the cost of assurance
Security controls succeed only if their operational cost is sustainable. An under-review interview study with 20 open-source maintainers examines how AI-generated vulnerability reports weaken the old link between polished presentation and technical validity. Maintainers responded by doing more independent verification and using AI for pre-triage while retaining human responsibility. The study identifies an attention asymmetry: generation becomes cheaper while verification remains expensive. It also shows how reporter-history heuristics can disadvantage newcomers. These findings keep my systems work grounded in human accountability, fair access, and usable evidence.
In 2025, I started an AI Assurance Programme within the Cybersecurity Group around three mutually reinforcing research thrusts spanning pre-deployment measurement, runtime security, and lifecycle governance. The programme follows a common principle: assurance should be based on explicit, reproducible evidence about what an AI system is expected to do, what authority it possesses, how it behaves, and how that evidence changes after deployment.
As explained below, Policy2Bench will translate human and institutional requirements into defensible evaluation programmes; agent interaction property graphs will make authority, provenance, capabilities, and execution visible in agentic systems; and AI governance will connect both forms of evidence to operational decisions.
5.1. Measurements that deserve trust
AI benchmarks often answer a precisely formulated question that is only loosely connected to the decision an organisation must make. A prospective adopter rarely begins with a benchmark name or a fully specified construct. The initial requirement is more likely to be "select a safe model for a multilingual public-service agent" or "identify the model that best supports our privacy obligations." Existing dynamic benchmarks can refresh test content, but they generally assume that the construct to be measured is already known.
Building on aiXamine, SSP-Bench, and Sayf-Eval, I will develop Policy2Bench, a compiler from vague model-selection requirements to auditable, grounded, dynamically generated LLM evaluations. It will transform a natural-language requirement into a typed representation and then select a sparse set of versioned policy atoms. Each atom will specify its scope, applicability, desired and prohibited behaviours, evidence requirements, test archetypes, scoring contract, provenance, and validity period. The compiler will use authoritative sources and explicit precedence rules to generate an evidence plan, synthesize source-grounded test items, independently derive their expected answers or rubrics, and freeze the complete inference, parsing, scoring, and aggregation pipeline.
An essential contribution will be to identify the boundary of behavioural evaluation. A model test may assess privacy-sensitive behaviour, abstention, or understanding of data-protection principles; it cannot establish an organisation's lawful basis, contractual controls, retention practices, or regulatory compliance. Policy2Bench will therefore produce two linked outputs: an executable benchmark for model-testable requirements and a separate evidence checklist for deployment, vendor, legal, and organisational claims. This prevents a behavioural score from being misrepresented as a compliance certificate (e.g., complying with GDPR regulations).
The scientific questions concern construct validity as much as generation quality. Can policy atoms faithfully capture the measurement semantics of existing benchmarks? Can an underspecified requirement be mapped to a smaller and more complete evaluation programme than retrieval or direct generation? Can the system recognise requirements that are not benchmarkable? Do dynamically generated suites remain attributable, contamination-resistant, discriminating, and stable enough to support a selection decision? Outputs will include an open policy library, requirement corpus, source registry, benchmark compiler, versioned evaluation specifications, and statistical methods for reporting uncertainty, trade-offs, and ranking stability.
5.2. Secure architectures for agentic AI
Agentic applications combine models, tools, prompts, memory, remote services, users, and delegated tasks. They should therefore be treated as dynamically composed security principals, not merely as LLM wrappers. Building on AgentSA's pre-deployment recovery of agent interaction graphs and authorisation rules, I will develop Agent Interaction Property Graphs (AIPGs): a framework-independent, temporal, and evidence-bearing representation of agentic software.
An AIPG will combine four complementary views. A deployment and identity plane will represent agents, users, workloads, hosts, models, credentials, and administrative ownership. A capability plane will over-approximate which agents, tools, resources, and external services are reachable. An authorisation plane will record why an action is permitted and how task-scoped authority is delegated or attenuated. An execution and provenance plane will record what actually occurred and which inputs, messages, agents, and tool invocations influenced an external effect. Graph edges will carry their source, collection method, confidence, freshness, signature, and administrative domain so that a policy can distinguish declared, observed, bilaterally confirmed, attested, and inferred facts.
The first stage will construct these graphs from source code, configuration, deployment manifests, identity policies, framework hooks, and telemetry from protocols such as MCP, A2A, and AG-UI. The second will address systems that cross organisational boundaries. Rather than disclosing complete internal topologies, independently administered applications will exchange signed, freshness-aware graph assertions or minimal path evidence sufficient to answer a security question. This creates research problems in attestation, revocation, Byzantine behaviour, topology privacy, selective disclosure, and the consistency of partial graph views.
The third stage, GraphGuard, will use independent reference monitors to enforce graph-grounded policies outside the LLM. Such policies can detect a path from attacker-controlled content to a privileged write, verify that authority narrows across delegation, identify confused-deputy behaviour, estimate the blast radius of a compromised agent, constrain cross-domain information flows, and detect execution edges absent from the declared architecture. Enforcement may deny or attenuate an action, require additional authorisation, redirect it to a safer tool, sandbox it, or quarantine a compromised component. The central research question is whether combining identity, authority, provenance, reachability, and execution enables protection against compositional attacks that remain invisible to prompt filters, tool allowlists, or local traces.
5.3. Governance as an operating system for assurance
The first two thrusts provide complementary evidence. Policy2Bench determines what should be evaluated before deployment and produces a versioned, policy-specific assurance profile. AIPGs and GraphGuard determine what authority and influence exist in the deployed system and enable runtime controls over consequential actions. I will connect them through AI governance: a lifecycle governance architecture spanning registration, pre-deployment certification, policy-as-code, local enforcement, privacy-preserving telemetry, incident response, and recertification.
The architecture will separate a governance control plane from integration and data planes. Organisations will retain custody of sensitive prompts, outputs, and operational data, while local gateways export only the evidence required for an oversight or certification decision. Policies compiled from human requirements will become traceable artefacts: some will produce behavioural tests, others will become runtime graph predicates, and still others will be routed to organisational evidence and human review. Runtime incidents and previously unseen interaction paths will, in turn, trigger new tests, policy revisions, or recertification rather than silently changing the evaluation target.
This creates a unified research agenda around evidence continuity. When does a model, prompt, tool, policy, or workflow change invalidate prior certification? How can a pre-deployment requirement be traced to a benchmark item, a runtime control, and an incident record? What is the minimum evidence that must cross organisational boundaries? How can logs be made tamper-evident without centralising private data? Blockchains will be used selectively where multiple parties require durable provenance, attribution, or incentive-compatible accounting, not as a default storage layer.
Closing Remarks. Across all three research thrusts above, I will study the human work of assurance: how experts resolve ambiguous requirements, how operators interpret warnings and contest automated decisions, and how verification costs affect maintainers and under-resourced communities. QCRI offers an exceptional environment in which to combine fundamental research with carefully scoped collaborations across industry and the public sector. The programme will publish reusable software, benchmark artefacts, policy representations, and evaluation corpora whenever ethics, confidentiality, licensing, and responsible disclosure permit. The intended outcome is an assurance stack whose measurements are valid, whose controls are enforceable, and whose governance mechanisms help people decide when an AI system deserves authority.
Yazan Boshmaf, Ildar Muslukhov, Konstantin Beznosov, and Matei Ripeanu
Yazan Boshmaf, Ildar Muslukhov, Konstantin Beznosov, and Matei Ripeanu
Yazan Boshmaf, Konstantin Beznosov, and Matei Ripeanu
Hootan Rashtian, Yazan Boshmaf, Pooya Jaferian, and Konstantin Beznosov
Yazan Boshmaf, Dionysios Logothetis, Georgos Siganos, Jorge Lería, Jose Lorenzo, Matei Ripeanu, and Konstantin Beznosov
Yazan Boshmaf, Dionysios Logothetis, Georgos Siganos, Jorge Lería, Jose Lorenzo, Matei Ripeanu, Konstantin Beznosov, and Hassan Halawa
Yazan Boshmaf, Husam Al Jawaheri, and Mashael Al Sabah
Yazan Boshmaf, Isuranga Perera, Udesh Kumarasinghe, Sajitha Liyanage, and Husam Al Jawaheri
Husam Al Jawaheri, Mashael Al Sabah, Yazan Boshmaf, and Aiman Erbad
Yazan Boshmaf, Charitha Elvitigala, Husam Al Jawaheri, Primal Wijesekera, and Mashael Al Sabah
Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf et al.
Khang Tran, Yazan Boshmaf, Issa Khalil, NhatHai Phan, Ting Yu, and Md Rizwan Parvez
Fatih Deniz, Yazan Boshmaf, Dorde Popovic, and Issa Khalil
Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, and Yazan Boshmaf