Start here

Home › AI Safety & Security

Hub · AI Safety & Security

AI Safety & Security: what is known

What the science says about AI risks, why chatbots make things up, how deepfake scams work, what prompt injection is, and simple privacy habits. Every claim links to a government, standards or research source.

Last reviewed Oct 3, 2026

100+independent experts wrote the 2026 International AI Safety Report
77%of vulnerabilities an AI agent found in one cyber competition (safety report)
12companies published or updated frontier safety frameworks in 2025
3risk categories in the report: malicious use, malfunctions and systemic risks

Last reviewed Oct 3, 2026. Safety and security explainers, not advice for any individual situation.

The short version

  • Chatbots state false things fluently. Even the makers say hallucinations are a fundamental challenge for all large language models.1
  • Fake voices, images and video are used in fraud now. The FBI has warned the public about it, and the International AI Safety Report says harms are well documented while data on how common they are is limited.2,3
  • AI tools that read untrusted content can be tricked. The UK’s NCSC says prompt injection may never be fully fixed the way older web flaws were, so risk has to be limited by design.4
  • What you type can be stored. The NCSC notes that a chatbot company can see the questions you ask it, store them, and will almost certainly use them to develop its service at some point.5 Settings differ by app: see Chat privacy.

Hallucinations

OpenAI defines hallucinations as “plausible but false statements generated by language models”. Its explanation: models first learn by predicting the next word in huge amounts of text, with no true/false labels attached, so arbitrary low-frequency facts (its example is a person’s birthday) cannot be predicted from patterns alone. It adds that most tests score only accuracy, which rewards guessing over saying “I don’t know”.1

OpenAI also gives an example from its own tests, below. It is a company result on one quiz, not a general error rate.1

SimpleQA result (OpenAI)Abstained (no answer)Right answerWrong answer
gpt-5-thinking-mini52%22%26%
OpenAI o4-mini1%24%75%

The International AI Safety Report 2026 says AI systems can generate non-existent citations, biographies and facts, that human oversight helps but invites over-reliance when answers are fluent and confident, and that human verification of outputs remains necessary in high-stakes settings such as medicine and finance.3

What to do. Treat names, dates, numbers, quotes and links as claims to check. If an AI gives a source, open it. Real examples where this went wrong in court and customer service are on How AI works. You can practice on the Reality check prompts.

What the science says is known, and not known

The International AI Safety Report 2026, written by over 100 independent experts with an advisory panel nominated by more than 30 countries and organizations, sorts risks into malicious use, malfunctions and systemic risks.3 A reading of its executive summary:

TopicWhat the report says is documentedWhat it says is uncertain
Scams, fraud and abusive contentAI systems are being misused to generate content for scams, fraud, blackmail and non-consensual intimate imagery.3Systematic data on how often it happens and how severe it is remains limited.
ManipulationIn experiments, AI-written content can be as effective as human-written content at changing beliefs.Real-world use for manipulation is documented but not yet widespread; it may increase.
CyberattacksAI can find software vulnerabilities and write malicious code; in one competition an AI agent found 77% of the vulnerabilities present. Criminal groups and state-associated attackers are using AI.3Whether attackers or defenders gain more from AI is uncertain.
ReliabilitySystems sometimes fabricate information, write flawed code and give misleading advice. Agents are riskier because they act with less human intervention.Current techniques reduce failure rates but not to the level many high-stakes uses need.
Loss of controlThe report says current systems lack the capabilities to pose such risks.They are improving in relevant areas, and it has become more common for models to tell test settings from real use, so some dangerous abilities could go undetected.
Open-weight modelsThey bring research and commercial benefits, especially for less-resourced groups.They cannot be recalled once released and their safeguards are easier to remove, which makes misuse harder to prevent and trace.

On what companies and governments are doing: 12 companies published or updated Frontier AI Safety Frameworks in 2025, most risk management remains voluntary, and a few jurisdictions are beginning to make some practices legal requirements. The report says technical safeguards are improving but that determined users can still sometimes get harmful outputs by rephrasing or splitting requests.3

Deepfakes and AI scams

The FBI’s December 2024 public alert describes how criminals use generative AI to make fraud more believable and to work at larger scale: text for fake profiles and fraud sites, images for fake IDs and profile photos, short cloned voice clips of a loved one in a made-up emergency, and video for real-time calls posing as executives or officials. It notes that creating or sharing synthetic content is not by itself illegal, but can be used for crimes such as fraud and extortion.2

The FBI’s listed protections, paraphrased:

  • Create a secret word or phrase with your family to verify identity.2
  • Hang up and call the person or organization back on a number you looked up yourself.2
  • Never send money, gift cards, cryptocurrency or other assets to people you have met only online or by phone, and do not share sensitive information with them.2
  • Limit public photos and voice recordings where you can, make social accounts private, and limit followers to people you know.2
  • Look for flaws in images and video (distorted hands, odd teeth or eyes, wrong shadows, lag, lip-voice mismatch), but the FBI itself says AI content can be hard to spot.2

For the FBI and FTC loss figures and the full list of habits, see AI scams and deepfakes. To report fraud to the FBI, use the IC3 site named in the alert.2

Prompt injection and AI agents

Prompt injection is when instructions hidden in content an AI reads (a web page, email or document) change what the AI does. OWASP ranks it first among its risks for large-language-model applications, and says research shows common fixes such as retrieval and fine-tuning do not fully prevent it.6

The NCSC explains why it is hard: inside a language model there is no separation between “data” and “instructions”, only the next token, so prompt injection may never be totally mitigated the way SQL injection can be. Its advice for builders is to reduce likelihood and impact, to use safeguards that do not rely on the model, and to give the AI no more privileges than the least-trusted party whose content it reads.4 OWASP likewise recommends least-privilege access and human approval for high-risk actions.6

What this means for you (our reading). The more an AI assistant can do for you (read your mail, browse, buy, send), the more a trick in a page or message it reads can matter. Check what an assistant is allowed to do, keep payment and account changes behind your own approval, and be wary of asking it to act on content from strangers.

Privacy tips

The NCSC’s guidance, written in 2023 but still on its site, says a language model does not automatically add what you type to its knowledge for other people to query. However, your query is visible to the company providing the service, is stored, and will almost certainly be used for developing the service or model at some point, and the company or its partners may be able to read it. It advises understanding the terms and privacy policy before asking sensitive questions.5

Who sets guidance

In the United States, NIST publishes a voluntary AI Risk Management Framework (January 2023) and a Generative AI Profile (NIST AI 600-1, July 26, 2024) intended to help organizations build trustworthiness into AI products.7 See the glossary entry on the NIST AI RMF. For laws, see AI rules and jobs and AI by Country.

Keep reading

What we left out. We did not include figures on the number of deepfake incidents or scam losses beyond those on the Scams page, because the safety report itself says systematic prevalence data is limited. We did not include claims about future, hypothetical AI risks beyond what the International AI Safety Report states, and we did not rank products by safety. See How we verify; report an error through the corrections page.

Sources for this page

  1. Why language models hallucinate. OpenAI.Published Sept 5, 2025 · checked Oct 3, 2026
  2. Criminals Use Generative Artificial Intelligence to Facilitate Financial Fraud (Alert I-120324-PSA). FBI Internet Crime Complaint Center (IC3).Published Dec 3, 2024 · checked Oct 3, 2026
  3. International AI Safety Report 2026. International AI Safety Report (chair: Yoshua Bengio; secretariat: UK AI Security Institute).Published Feb 3, 2026 · checked Oct 3, 2026
  4. Prompt injection is not SQL injection (it may be worse). UK National Cyber Security Centre.Published Dec 8, 2025 · checked Oct 3, 2026
  5. ChatGPT and large language models: what’s the risk?. UK National Cyber Security Centre.Published Mar 14, 2023 · checked Oct 3, 2026
  6. LLM01:2025 Prompt Injection. OWASP Gen AI Security Project.Current page · checked Oct 3, 2026
  7. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology.Published July 26, 2024 · checked Oct 3, 2026