AI vendor due diligence for social service agencies: turn assurances into evidence
A vendor questionnaire is useful when every broad claim leads to a clear scope, evidence you can inspect, a commitment you can rely on and a test your agency can run.
In brief
Translate each vendor assurance into five things: its exact scope, documentary evidence, a contractual commitment where appropriate, a named owner and a way for the agency to verify the claim or leave the service.
Introduction
An AI vendor can answer a long questionnaire and still leave the agency unsure what happens to its information. The familiar phrases sound reassuring: enterprise-grade, encrypted, government-grade, no training on customer data, compliant. Each may describe something real. Each may also cover a much narrower point than the team assumes.
Good due diligence makes the claim answerable. Which product, plan, feature, data type, location and period does it cover? What document or test result supports it? Which part belongs in the contract? Who owns the control after launch? How will the agency check that it works for its own workflow?
This resource sits alongside the broader responsible AI governance guide for social service agencies. If the immediate question is whether staff may submit a document to a generative AI service, start with Can I paste this into ChatGPT?.
Start with the proposed use
A product cannot be assessed in the abstract. Write a one-page use statement before sending questions to vendors:
- the staff and service users involved;
- the task, source material and expected output;
- every type of information the tool would receive or infer;
- the action someone may take from the output;
- the likely effect of a serious error, disclosure or outage; and
- the existing process that remains available if the tool fails.
The same product may be reasonable for drafting a public event description and unsuitable for summarising identifiable case records. Product controls also differ by plan, account configuration and enabled feature. Ask the vendor to confirm its answer against the written use statement rather than a generic “nonprofit use”.
PDPC's 2024 AI guidance says service providers can support customers, while the organisation choosing the system bears primary responsibility for ensuring that it can meet the organisation's PDPA obligations.[2] The 2026 generative AI guidance also describes responsibilities across model providers, system providers and system deployers, including accountability, protection, retention and purpose limitation.[1] Legal and data-protection reviewers should apply that guidance to the actual arrangement.
Keep three kinds of assurance separate
A useful evidence record labels what each item can establish.
Vendor evidence describes a current control or past activity. Examples include a data-flow diagram, subprocessor list, security test summary, incident procedure, deletion specification or completed assurance report. Check its scope, date, author and exceptions. A polished policy is weaker evidence for a technical claim than a current system-specific artefact.
Contract terms allocate duties and remedies. They can cover permitted processing, notice of changes, retention, incident notification, audit information, subprocessor controls, service levels, export, deletion and transition support. A contract does not prove that a control works today. It gives the agency an enforceable commitment and a route to act if the commitment is missed.
Agency testing examines the real workflow under realistic conditions. It can reveal failures that a general benchmark never touched: Singlish or code-switching, local scheme names, scanned forms, unusual document layouts, long context, weak connectivity, staff permission errors and outputs that sound certain when the source is uncertain. Testing does not replace contractual protection or security assurance.
Keep all three in the same decision record. Record gaps plainly, along with the person accepting, resolving or stopping over each gap.
Ten due-diligence conversations
1. What purpose and user group is the product designed for?
Ask for the supported workflow, intended users, known exclusions and the conditions under which the vendor's claims were tested. Compare that scope with the agency's one-page use statement.
2. Where does every copy of the information travel?
Map prompts, files, audio, outputs, metadata, logs, caches, backups, support records, connected apps, model APIs and analytics. Include storage and processing locations, even when a copy is short-lived.
3. What secondary uses occur?
Separate model training and fine-tuning from evaluation, abuse monitoring, product improvement, human review and analytics. Ask how the answer changes by plan, setting, feature and opt-out choice.
4. How are retention and deletion handled?
Request separate periods and deletion triggers for each artefact. Ask what “deleted” means for active systems, caches, backups and subprocessors; who initiates it; and what evidence the agency receives.
5. Who can access the service and its data?
Cover agency roles, vendor staff, privileged administrators, subprocessors and support access. Ask about least privilege, approval, logging, periodic review and notification when the subprocessor list changes.
6. How is the system secured?
Ask about secure development, tenant separation, encryption and key management, vulnerability handling, audit logs, recovery and incident response. Send material claims to a security reviewer rather than relying on procurement staff to interpret them alone.
7. What has been tested, and what remains for us to test?
Request methods, versions, dates, datasets, thresholds, results, limitations and remediation status. Build an agency test set around the actual language mix, documents, edge cases and unacceptable failures.
8. Where does human control sit?
Check permissions, draft states, review queues, source access, audit trails, escalation, override and fallback. Confirm that an unreviewed output cannot quietly enter a case, donor or service system.
9. How will change be managed?
Ask for notice of model, feature, term, price, data-location and subprocessor changes. Define which changes require agency review, retesting, reconfiguration or a pause.
10. Can the agency leave cleanly?
Confirm usable export formats, migration help, timing, costs, retained records, deletion evidence and continuity if the service or vendor ends. Test an export before the agency depends on the product.
Turn a broad assurance into a decision record
| Vendor assurance | Evidence to request | Contract or operational follow-up |
|---|---|---|
| “We do not train on your data.” | A dated description of covered data, products, plans, settings and exceptions, including evaluation, monitoring and human access. | State permitted uses and change notice in the agreement. Verify account settings and connected services before real data is allowed. |
| “Data is encrypted.” | Architecture showing encryption in transit and at rest, the covered stores and backups, key ownership and management, and any points where data is decrypted. | Record the required controls and incident duties. Have a security reviewer check the architecture and configuration evidence. |
| “We retain data only as needed.” | A retention schedule for each artefact, deletion workflow, backup treatment and subprocessor propagation. | Set applicable periods, deletion events and evidence in the agreement. Run a deletion exercise with test data. |
| “Our model is accurate.” | Test design, model and application version, datasets, metrics, thresholds, results, limitations and known failure groups. | Agree acceptance and stopping rules. Test the agency's documents, languages and consequential errors before live use. |
| “We are enterprise-grade.” | Specific controls for identity, permissions, audit, availability, support, recovery, vulnerability handling and tenant isolation. | Choose the controls needed for this use, include service and incident terms where appropriate, then verify configuration and recovery steps. |
| “We use approved subprocessors.” | Current names, roles, data received, processing locations, assurance information and the vendor's review process. | Define notice, objection or exit rights for material changes. Keep an owner and review date for the list. |
Read security evidence in context
CSA's Guidelines on Securing AI Systems recommend secure-by-design and secure-by-default practices across the AI lifecycle. They cover familiar cybersecurity risks, including supply-chain attacks, alongside AI-specific threats such as adversarial machine learning. The companion guide is explicitly non-prescriptive and maintained as a living resource.[3]
Ask the security reviewer to match evidence to the proposed architecture. A certificate, report or penetration-test summary may be useful, but its title alone says little. Check the system boundary, services excluded, test date, unresolved findings and whether the agency will use the same hosting, identity and integration pattern.
The questionnaire should also cover incident operations. Who tells the agency what happened, through which channel, within what agreed time? What logs can be preserved? Can the vendor identify affected tenants and artefacts? Who coordinates with subprocessors? What service can continue safely while investigation is under way? The final wording belongs in the actual agreement and incident plan.
Test the use the agency will run
AI Verify provides a testing framework spanning governance principles including security, robustness, fairness, data governance, accountability and human oversight. Its process checks can be supported by documentary evidence.[5] Alignment or a report can contribute to due diligence. It does not certify that a vendor is safe, PDPA compliant or suitable for every social-service workflow.
Build a small agency test pack with synthetic or properly approved material. Include routine examples, difficult examples and failures the team refuses to accept. For a document assistant, that might include crossed-out text, poor scans, conflicting dates, local acronyms, bilingual passages, uncertainty, allegations and blank fields. Test the whole application, including retrieval, integrations, permissions and review screens, rather than sending isolated prompts to the model.
Agree thresholds before the demonstration. Record the version, settings, test inputs, expected result, observed result and reviewer decision. Keep the existing workflow available. A vendor can help reproduce an issue, but the agency decides whether the result is acceptable for the intended service.
NCSS's Tech-and-GO! page provides separate toolkits for digital strategy, technical evaluation and project implementation.[4] That separation is useful: selecting a plausible solution is followed by careful implementation, ownership and monitoring. It is not an endorsement of this checklist or any product.
Treat missing evidence as an open decision
A vendor may decline to share a detailed report because it contains security-sensitive information. A small supplier may have sound controls without a familiar assurance package. A mature provider may answer only through a standard portal. Record what was requested, what was supplied, what could be reviewed under confidentiality and what remains unknown.
Then choose a disposition for each gap: obtain another artefact, narrow the use, add a contract term, test a compensating control, accept the uncertainty through the agency's authorised process, or stop. The owner and reasoning matter more than a coloured score. One unresolved issue involving sensitive case data can outweigh a folder full of polished documents.
Avoid turning the questionnaire into a product ranking. Two agencies can reach different decisions about the same service because their information, integrations, professional duties, fallback options and consequences differ. Even within one agency, approval for public communications does not extend automatically to case work. The evidence record should state the approved boundary in plain words.
Evidence gates before commitment, pilot and live use
| Gate | Minimum record before proceeding | Likely owners | Decision |
|---|---|---|---|
| Before contract | Written use statement; initial data-flow map; vendor evidence register; key gaps; security and data-protection review; agreed change, incident, subprocessor, export and deletion terms. | Project owner, procurement, DPO/legal, IT/security | Contract, narrow the use, resolve a gap, choose another route or stop. |
| Before pilot | Approved test data; configured roles and settings; test plan and thresholds; fallback; incident route; named reviewers; confirmation that prohibited data cannot enter the pilot. | Project owner, service lead, DPO, security, vendor technical contact | Start a bounded pilot, redesign the test or wait. |
| Before live use | Agency test results; unresolved failures and acceptance owner; staff guidance; access and audit checks; retention and deletion operation; export test; monitoring and review date. | Accountable executive, service owner, DPO, security, operations | Approve the defined use, restrict it, extend the pilot or stop. |
Make the record usable after selection
Due diligence ages quickly. Keep a short register with the product and plan, approved use, data classes, evidence dates, contract commitments, settings, subprocessors, agency test version, open issues, owners and next review. Attach the one-page use statement so future staff can see the boundary that was assessed.
Set review triggers as well as a calendar date. A new model, connector, agentic feature, data location, subprocessor, term or use case can change the answer. So can an incident, a serious output failure or evidence that staff are bypassing review.
Plan the exit while the relationship is healthy. Export representative synthetic records and check that another system or ordinary process can use them. Record what will remain after termination, for how long and why. Ask how deletion is evidenced. The Commissioner of Charities' data-protection guide discusses due diligence, cloud storage and jurisdiction, and written outsourcing arrangements for outsourced electronic personal-data storage.[6] An agency should check its currency and fit with legal and DPO advisers before relying on it.
No checklist makes a vendor “approved”. The decision applies to a defined use under stated conditions. Clear boundaries make it easier to proceed carefully, and easier to pause when evidence no longer supports the use. This resource provides general information rather than legal, procurement or security advice.
Sources
- [1] Personal Data Protection Commission, Advisory Guidelines on Use of Personal Data in Generative AI, issued 20 July 2026
- [2] Personal Data Protection Commission, Advisory Guidelines on Use of Personal Data in AI Recommendation and Decision Systems, issued 1 March 2024
- [3] Cyber Security Agency of Singapore, Guidelines and Companion Guide on Securing AI Systems, 15 October 2024
- [4] National Council of Social Service, Tech-and-GO! consultancy guides, updated 6 February 2025
- [5] AI Verify Foundation, AI Verify Testing Framework, updated for generative AI in May 2025
- [6] Commissioner of Charities, Data Protection Guide for Charities: Managing & Securing Electronic Personal Data
Check currency and application with legal/DPO advisers before agency reliance.
About the author
Darren writes for Social Tech Guild about practical, responsible uses of technology in Singapore's social service sector.
Review the use and evidence before selecting the tool
Social Tech Guild can help your project, data and service leads turn a proposed workflow into a vendor-neutral question set and evidence record.
Discuss a vendor-neutral reviewCould a small tool make your team’s work lighter?
Tell me about a task that keeps taking time. We can look at it together and see whether a small volunteer project could help.
Please do not include client-identifying information.