Vision Nexera

Saudi PDPL Compliance for AI Systems: Chatbots, Voice Agents & RAG

September 29, 2026 · Asma Nawaz · Technical Content Writer, Vision Nexera · 1 min read

A build guide to Saudi PDPL compliance for AI. How to make chatbots, voice agents, and RAG pipelines pass SDAIA scrutiny, with the exact controls for consent, residency, and data subject rights.

Saudi PDPL Compliance for AI Systems: How to Build Chatbots, Voice Agents, and RAG Pipelines That Pass SDAIA Scrutiny

Three kinds of AI system now show up in almost every Saudi deployment: a chatbot on the website or WhatsApp, a voice agent on the phone lines, and a RAG pipeline answering questions from the company's own documents. All three are built to consume personal data, and under the Personal Data Protection Law that makes each of them a surface the Saudi Data and Artificial Intelligence Authority can inspect. The good news is that passing that inspection is an engineering problem with known answers. This is a build guide to those answers: the specific controls that make a chatbot, a voice agent, and a RAG pipeline defensible under PDPL and aligned with SDAIA's AI expectations.

It is a companion to our data residency guide for PDPL compliant AI in Saudi Arabia, which covers where Saudi data and models can legally run and how to choose a partner. That piece is the foundation. This one goes a level deeper into the obligations themselves and how they translate into architecture for each system type, so it is worth reading that one first if you have not.

One thing up front, and it matters. This is engineering guidance, not legal advice. PDPL is a binding law with an active regulator, and any real deployment should be reviewed with Saudi counsel and, where required, registered with SDAIA. What follows is how to build systems that support compliance, not a replacement for a lawyer.

What SDAIA scrutiny actually covers

It helps to know that scrutiny comes in two layers. The first is the binding one: the Personal Data Protection Law and its Implementing Regulations, enforced by SDAIA, in force since September 2023 with full enforcement since September 2024. By early 2026 SDAIA had issued dozens of enforcement decisions, with administrative penalties up to SAR 5 million and criminal sanctions for the intentional disclosure of sensitive data. This is the layer that decides whether your system is legal.

The second is SDAIA's AI specific guidance: its AI Ethics Principles and its Generative AI Guidelines. These are largely soft law rather than binding statute, but they set the expectations SDAIA judges responsible AI against, and they are explicit about the things that matter most for chatbots, voice agents, and RAG. They call for transparency, human oversight, explainability, fairness, privacy, and active management of risks like hallucinations, bias, and unreliable content, and they sort AI systems into risk tiers from little or no risk up to unacceptable. A system that handles personal data through a language model sits squarely in the range these guidelines are written for.

Put simply, to pass SDAIA scrutiny an AI system has to do two things at once: satisfy the hard data protection rules of PDPL, and behave the way SDAIA's AI guidance says a responsible system should. The rest of this guide is about building for both.

The compliance foundation every AI system needs

Before the system specific parts, there is a set of controls that a chatbot, a voice agent, and a RAG pipeline all need. Get these right once and most of the work carries across all three.

  1. Classify the data first. Apply the National Data Management Office data classification tiers to everything the system will touch before designing anything. The classification decides what must stay in the Kingdom and which controls apply. Remember that under PDPL the sensitive category is broad: it includes biometric, health, racial or ethnic, religious, ideological and political data, and, unusually, national ID numbers and credit information. A system that collects a national ID is handling sensitive data.

  2. Establish a lawful basis and capture consent properly. Consent is the default basis, and PDPL sets a high bar: it must be explicit, specific to each purpose, documented in a way you can later verify, and revocable at any time. For sensitive data, including national IDs, the bar is higher still, explicit consent with detailed disclosure, and no other lawful basis can substitute. Consent can be written, verbal, or electronic, but in every case the time and means must be recorded.

  3. Show a proper privacy notice. Before or at the moment of collection, tell the person who the controller is, what you collect, why, on what basis, who receives it, whether it leaves the Kingdom, how long you keep it, and how to exercise their rights, in clear language an ordinary person can understand.

  4. Minimize and limit purpose. Collect only what the task needs, and use it only for the purpose you disclosed. Repurposing personal data quietly, for training or analytics, without a basis is one of the fastest routes to an enforcement problem.

  5. Keep data and inference in the Kingdom. Host the application, database, and logs in a Saudi region, and run the model in region or self hosted so personal data in prompts does not cross the border. The residency guide covers the cloud options in detail. Encrypt data at rest and in transit, and keep the keys under your control inside the Kingdom.

  6. Control access and log everything. Enforce least privilege access to data and tools, and keep an audit trail of access and actions, both to contain the system and to demonstrate accountability to SDAIA.

  7. Support data subject rights. Individuals have the right to be informed and to access, correct, delete, and port their data, and to withdraw consent. Build the machinery to honor these requests, generally within thirty days, extendable once by up to thirty more.

  8. Set retention limits and deletion. Decide how long each category of data lives, delete it when the purpose ends, and make sure deletion actually reaches every store the data landed in.

  9. Prepare for breaches. Have an incident response plan that can notify SDAIA within seventy two hours of becoming aware of a breach that risks harm, with the prescribed details, and notify affected individuals without undue delay where the risk is serious.

  10. Keep records and, where required, a DPO. Maintain a written record of processing activities covering purposes, categories, retention, recipients and transfers, and security measures, appoint a data protection officer if you fall in scope, and register with SDAIA on the National Data Governance Platform.

  11. Keep a human accountable and the system explainable. SDAIA's AI guidance expects human oversight of consequential decisions and a system that can explain how it reached an outcome. Design for a human checkpoint and for traceability from the start.

With that foundation in place, here is what changes for each of the three system types.

Building a PDPL-compliant chatbot

A chatbot looks harmless and is quietly one of the leakiest systems you can deploy. It invites people to type freely, which means it collects names, phone numbers, order details, and, whether you asked for them or not, national IDs, health complaints, and other sensitive data dropped into a text box. Every message it forwards to a language model is a potential cross border transfer, every conversation it stores is a personal data record, and the temptation to train on those conversations is a purpose limitation problem waiting to happen.

Start the conversation with consent and a notice, not with a question. Before the bot collects anything, it should state who is behind it, that it is an automated system, what it will do with what the person types, and how they can opt out or reach a human, with a clear link to the full privacy notice. In Saudi Arabia this needs to work in Arabic, and the consent needs to be logged with a timestamp so you can prove it later.

Minimize what it asks for, and redact what it does not need. Design the flows to request only the fields the task requires, and run personal data detection on the input so that identifiers are stripped or masked before the text is stored or sent to a model that sits outside the Kingdom. Edge redaction here removes a whole class of transfer and sensitivity problems cheaply. If the bot genuinely needs a national ID, treat that as sensitive data with heightened explicit consent, and never let it be collected casually mid chat.

Keep the model and the transcripts inside the Kingdom. Run inference through an in region endpoint or a self hosted model so message content does not leave, store conversations in a Saudi region with encryption and access control, and do not use customer conversations to train or fine tune a model without a lawful basis to do so. Set a retention period for transcripts and delete on schedule.

Make data subject rights real. A person should be able to ask what the bot holds about them, correct it, and have it deleted, which means you need to be able to find and remove a specific individual's conversations on request. Guard against the bot itself becoming a leak: it should refuse to reveal one user's data to another, and it should not coax people into sharing sensitive information it has no need for. Because chatbots are conversational agents, the same discipline we apply to any AI agent applies here, and the safe handling of tool and database access is the same pattern described in our natural language to database MCP server, where the model's output is treated as untrusted and destructive actions are blocked at a layer it cannot talk its way past.

Building a PDPL-compliant voice agent

A voice agent inherits every chatbot concern and adds two that are specific to voice, and both are serious. First, the call recording and its transcript are personal data, and they often contain far more than a text chat because people say things out loud that they would never type. Second, a voice print is biometric data, which PDPL treats as sensitive, so the moment you store or analyze the characteristics of someone's voice you are in the heightened consent and control regime.

Capture consent out loud, at the start of the call. Before recording or processing begins, the agent should disclose that the call is with an automated system and may be recorded and processed, state the purpose, and give the caller a way to decline or reach a person. PDPL accepts verbal consent, but it must be documented and verifiable, so the disclosure and the caller's response need to be captured and stored. Doing this in clear Arabic is part of doing it properly.

Keep speech to text and text to speech in the Kingdom, or redact before they leave. The transcription and voice synthesis services are where voice data most often crosses the border, because many providers route audio to servers abroad by default. Use providers that can run in region or under a contract that keeps the data in the Kingdom, or redact identifiers from transcripts before any offshore step. The real time nature of a call makes this harder than in a chatbot, which is exactly why it needs to be designed in rather than bolted on. The engineering of the voice stack itself is covered in our guide to AI voice agent development.

Treat the voice print as sensitive, and prefer not to keep it. Unless voice biometrics are genuinely central to your use case, the cleanest path is to avoid storing voiceprints at all. If you do need them, for authentication say, treat them as sensitive data with explicit consent and the strictest controls, and be clear that you cannot repurpose them for anything the caller did not agree to. Keep recordings only as long as the purpose requires, store them encrypted, restrict access tightly, and make sure a caller can exercise their rights over a recording, including having it deleted. And keep a clean path to a human, both because callers are entitled to one and because SDAIA's guidance expects human oversight.

Building a PDPL-compliant RAG pipeline

RAG is different in character from the other two. A chatbot and a voice agent mostly collect personal data from the person in front of them. A RAG pipeline points a language model at a body of your own documents, and those documents are very often full of other people's personal data: HR files, customer records, contracts, support histories, medical or financial documents. That changes the risk profile in three specific ways, and each has a control.

The corpus itself is the exposure, so classify and place it carefully. Before you index anything, classify the document set against the NDMO tiers, because the pipeline is only as compliant as the most sensitive thing in it. Host the vector database, the ingestion pipeline, and the query logs inside the Kingdom, not just the application, since personal data in embeddings and in retrieved chunks is still personal data. Where a document's personal identifiers are not needed for retrieval to work, redact or pseudonymize them before they are embedded.

Retrieval must respect permissions, or it becomes an unauthorized disclosure engine. This is the control most naive RAG systems miss and the one most likely to draw SDAIA's attention. If retrieval simply returns the most relevant chunks regardless of who is asking, it will happily surface a document one employee was never allowed to see to another who asks the right question. Retrieval has to be filtered by the requesting user's permissions, so the system can only ever return what that person is entitled to access. Under PDPL, leaking a personal data document across a permission boundary is exactly the kind of unauthorized disclosure that carries penalties.

Deletion has to reach the vector store, and answers have to be traceable. When a source document is deleted, or when a person exercises their right to have their data erased, that deletion must propagate to the embeddings in the vector database, not just to the original file, or you have kept a copy of data you were told to delete. Keep inference in region or self hosted so the retrieved personal data in each prompt does not cross the border, and build in citations so every answer can be traced back to its source, which serves both SDAIA's expectation of explainability and your own ability to audit what the system said and why. The underlying design decisions, and when retrieval is even the right tool versus fine tuning, are covered in RAG versus fine tuning and in our RAG and LLM application development work.

When a model or service only exists offshore

Sometimes the model or a component you need genuinely has no in Kingdom option. PDPL does not forbid this outright, but it turns the flow into a regulated cross border transfer, which needs a valid basis and safeguards. Because SDAIA has not published a list of adequate countries, that means SDAIA approved Standard Contractual Clauses or Binding Corporate Rules, a documented transfer risk assessment, transfers limited to the minimum data required, and nothing that prejudices the Kingdom's interests. The practical order of preference is simple: keep it in the Kingdom, and if you cannot, redact the personal data out before it leaves, and if you cannot do that either, transfer it under the proper safeguards and document everything. The residency guide goes through this in more detail.

The SDAIA scrutiny checklist

Whichever of the three you are building, this is the consolidated list to hold it against before it goes live:

  1. Data classified against NDMO tiers, with sensitive data, including national IDs and biometrics, identified.

  2. A lawful basis for every processing purpose, with explicit, specific, documented, revocable consent where consent is the basis, and heightened consent for sensitive data.

  3. A clear Arabic privacy notice shown before or at collection, with all the required disclosures.

  4. Data minimization and purpose limitation enforced in the design, not just the policy.

  5. Application, data stores, logs, and model inference kept in the Kingdom, with offshore flows only under proper safeguards.

  6. Encryption at rest and in transit, with keys held in the Kingdom.

  7. Least privilege access, and for RAG, permission filtered retrieval so nothing leaks across boundaries.

  8. Data subject rights supported end to end, including deletion that reaches the vector store, within the required timelines.

  9. Retention limits set and deletion actually performed.

  10. A breach response plan that can notify SDAIA within seventy two hours.

  11. A written record of processing activities, a DPO where required, and registration with SDAIA.

  12. Human oversight of consequential decisions, traceable and explainable outputs, and active handling of hallucination and bias risk.

How Vision Nexera builds for SDAIA scrutiny

Vision Nexera designs chatbots, voice agents, and RAG pipelines with these controls built in from the first architecture decision rather than retrofitted before an audit. In practice that means classifying the data before designing the system, deploying the application and its data stores in a Saudi region, running inference in region or on a self hosted model so personal data stays in the Kingdom, redacting personal data at the edge where a component has to sit offshore, filtering retrieval by permission so a RAG system never leaks across boundaries, holding encryption keys in the Kingdom, and building consent capture, data subject rights, audit logging, and a human checkpoint into the system rather than around it. The company grounds its systems through RAG and LLM application development, connects them to existing tools through AI integration behind a clean boundary, and publishes its security and data practices openly. Its presence in Doha gives Gulf clients real proximity, visible on the locations page, and the way it runs an engagement, including the data and residency decisions, is documented in its process.

The limitation is the same one worth repeating on any compliance topic. Vision Nexera builds the systems that support PDPL compliance. It is not a law firm, and under the law the client remains the data controller and carries the compliance responsibility. The right setup pairs a compliance first build with Saudi legal counsel and registration with SDAIA. Any partner who says their software alone makes you compliant is describing something the law does not offer.

Frequently asked questions

What does SDAIA actually scrutinize in an AI system?

Two layers. The binding one is the Personal Data Protection Law: lawful basis and consent, data minimization, cross border transfer safeguards, data subject rights, breach notification, records of processing, and security. The second is SDAIA's AI guidance, its AI Ethics Principles and Generative AI Guidelines, which expect transparency, human oversight, explainability, and active management of risks like hallucinations and bias. A system needs to satisfy both.

Can I use ChatGPT or Claude for a Saudi chatbot or voice agent?

You can use those models, but their direct APIs route data outside the Kingdom by default, which makes each call a cross border transfer of any personal data in the prompt. To keep it compliant, run the model through an in region endpoint, self host an open weight model, or redact personal data before the call and put the proper transfer safeguards in place.

Is a voice print sensitive data under PDPL?

Yes. Biometric data, which includes a voice print, is sensitive personal data under PDPL, so it requires explicit consent with detailed disclosure and the strictest controls, and it cannot be repurposed. Unless voice biometrics are central to your use case, the cleanest approach is not to store voiceprints at all.

Are national ID numbers treated as sensitive data?

Yes, and this catches many teams out. PDPL treats national ID numbers and credit information as sensitive, alongside the usual categories like health and biometrics. A chatbot or voice agent that collects a national ID is handling sensitive data and needs heightened explicit consent, with no substitute lawful basis.

If someone asks us to delete their data, does that include the vector database?

Yes. A deletion request has to reach every place the data lives, including the embeddings in your RAG system's vector store. Deleting the original document while leaving its embedded copy in the vector database means you have not actually deleted the data, which is a compliance failure.

What is the breach notification timeline?

You must notify SDAIA within seventy two hours of becoming aware of a personal data breach that risks harm to data subjects, with the prescribed details, and notify affected individuals without undue delay where the risk to them is serious. This requires a tested incident response process, not an improvised one.

Do I need a data protection officer?

Certain controllers are required to appoint a data protection officer under PDPL. Whether you fall in scope depends on your processing, so this is a question for your counsel, but the safe assumption for an organization running AI systems on personal data at any scale is that formal accountability, records, and SDAIA registration will be expected.

Is this article legal advice?

No. It is engineering guidance on building AI systems that support PDPL compliance and align with SDAIA's AI expectations. Any real deployment should be reviewed with Saudi legal counsel and registered with SDAIA where the law requires it.

Build it compliant, or rebuild it later

The pattern across chatbots, voice agents, and RAG pipelines is the same. The compliance work is not a layer you add at the end, it is a set of architecture decisions, classify the data, keep it and the model in the Kingdom, capture consent properly, control access, honor rights including deletion, and keep a human accountable, that are cheap to design in and expensive to retrofit. SDAIA is enforcing, the penalties are real, and the systems that pass scrutiny are the ones that were built for it from the first sprint.

If you are building a chatbot, a voice agent, or a RAG pipeline for the Saudi market, start with a written scope that puts compliance and residency first. Book a scoping call with Vision Nexera and leave with an architecture designed to pass SDAIA scrutiny, an honest estimate, and a clear view of what compliance will take.

Related reading

PDPL-Compliant AI Agent Development in Saudi Arabia, a data residency guide: https://www.visionnexera.com/insights/pdpl-compliant-ai-agent-development-saudi-arabia

Top AI Voice Agent Development Companies in Pakistan 2026: https://www.visionnexera.com/insights/top-ai-voice-agent-companies-pakistan-2026

RAG vs fine-tuning, a decision framework that fits on one page: https://www.visionnexera.com/insights/rag-vs-fine-tuning

Anatomy of a natural language to MongoDB MCP server: https://www.visionnexera.com/insights/mongodb-mcp-server-architecture

AI Agents service: https://www.visionnexera.com/services/ai-agents

RAG and LLM Applications service: https://www.visionnexera.com/services/rag-development

AI Integration service: https://www.visionnexera.com/services/ai-integration

Security and data practices: https://www.visionnexera.com/security


Next step

Tell us what you're building.

A 30-minute scoping call gets you a written scope and an honest estimate, including whether AI is even the right tool for it.

Prefer async? hello@visionnexera.com · We reply within one business day.

ASKArchitect⌘K