Responsible AIfor Behavioral Health
Framework · For clinicians and leaders

Behavioral Health Responsible AI Framework

A practical five-domain guide to useful tools, clear decisions, and better support for care teams.

By Cody Saunders, LMSW · October 5, 2026

AI can help care teams spend less time on busywork. It can help draft notes, explain services, and organize office work. The aim is better care, stronger bonds with clients, and a workday staff can manage. A good plan turns that aim into a clear task, a fair test, and a shared decision.

This framework is a guide for that plan. It has five domains, or areas to review: Clinical Appropriateness, Privacy & Security, Reliability & Safety, Human Oversight, and Governance & Accountability. Use all five for each use of AI. A tool may fit one task and fail to fit another.

This is an evolving learning and planning framework. It has not been validated as a test. It is not a certification, a legal opinion, or proof that a tool is safe. The proposed work steps below are Cody's practical approach. They draw on established sources but do not replace laws, professional standards, or your team's policies.

A shared plan for useful AI.

Care team members making a plan together.

Start with the work people need help doing

Name a problem in plain words. “Our therapists spend too long fixing visit notes” is useful. “We need AI” is too broad. Write down how the work is done now and what a better day would look like. Include clients and staff who will be affected.

Keep clinical and office tasks distinct. A draft staff agenda is office work. A draft client note supports care and may also affect billing. A tool that suggests treatment steps has a more direct role in care. Even office work can involve private health data. The task name alone does not tell you how much review it needs.

Before testing, make a short use record. Include the task, tool and version, data types, users, tool owner, final reviewer, and approval date. State what the tool may do and what it may not do. Then work through the five areas below.

1. Clinical Appropriateness: fit the tool to the care task

The goal: use a tool where its purpose and evidence fit the people and work involved.

Start with a narrow scope. An approved note tool might organize facts from a visit. That does not mean it may choose a diagnosis or decide the next treatment. State that boundary in the use record and staff training.

Review evidence for the exact task. A study of office notes in primary care does not prove better therapy outcomes. A seller's demo shows what the seller chose to show. Label seller claims as seller claims. Look for independent studies when they exist. Record what is known and what remains untested in your setting.

Consider the people served. Test plain language, reading needs, language access, age, and the service setting. A standard draft may miss what matters to a client. A qualified professional should judge whether client-facing content fits the care plan. Check that AI use supports the relationship rather than drawing attention away from it.

Put it into practice: define the task and limits; review relevant evidence; test examples that reflect the service; record who judged the fit. Include a route for clients or staff to raise concerns.

Keep as proof: the written scope, evidence notes, sample results, reviewer comments, and any limits placed on use.

Hold the use if no qualified person can judge clinical fit, or if the proposed task goes beyond the evidence and controls the team can support. Seek a narrower use or stronger evidence before moving ahead.

For social workers, competence includes the knowledge and skills needed to use technology in practice. This is professional ethics guidance, not a universal AI law. NASW Code of Ethics, section 1.04.

2. Privacy & Security: know and control the data path

The goal: send only allowed data through an approved, understood path.

Map what goes in, where it is processed, who can access it, and where results go. Include recordings, requests, logs, support access, and firms hired by the seller. Record storage periods, deletion terms, and any use of data for model training. A setting that turns off training does not settle every privacy issue.

Use the existing AI data-handling worksheet for this work. Do not put real client details into a planning worksheet. Use data types or made-up examples.

For a HIPAA-covered entity, a cloud provider that handles electronic protected health information on its behalf is generally a business associate. A HIPAA-compliant business associate agreement and compliance with the other applicable HIPAA duties are needed. An agreement alone does not approve every data use. HHS also calls for risk analysis and risk management. HHS cloud guidance.

Some substance use disorder records fall under 42 CFR Part 2, a separate federal rule. It does not cover every record that mentions substance use. HHS states that compliance with its 2024 final rule was required by February 16, 2026. Have the privacy team decide what applies to the records and proposed use. HHS Part 2 guidance.

Also check state rules, consent needs, special record protections, and contract terms with the right advisers. Staff need a clear allowed-data rule, not a guess based on a tool's “secure” label.

Put it into practice: complete the data map; have privacy and security staff review the exact service and contract; set access controls; confirm needed approvals before using live data.

Keep as proof: the data map, agreements, access settings, review findings, and written approval for the specific data use.

Hold the use when a required agreement is missing, the data path is unclear, or the team cannot confirm that the disclosure is allowed. Start with made-up test data while the team resolves the gap.

3. Reliability & Safety: test what matters in real work

The goal: know how the tool performs, and have a workable response when it does not perform well.

Build test cases before a live pilot. Include clear facts, missing facts, unclear wording, and cases outside the intended scope. Use made-up cases at first. For a note tool, check whether the draft changes meaning, adds a finding, or loses the client's words. Count errors by type. A small typo and an invented clinical claim should not count as the same issue.

Choose test standards before looking at the results. Define what counts as a usable draft and what calls for a pause. There is no single pass rate in this framework. The standard must fit the task and its possible effects. Do not average away a serious error because many easy cases worked.

Check performance across the groups and languages the service is meant to support. Record gaps in the test set. A successful small pilot does not prove performance for everyone or long-term benefit.

Test the whole workflow. Include outages, rejected drafts, corrections, and the usual process used when the tool is unavailable. For tools that read outside files or messages, security staff should test how the tool handles content that tries to change its instructions. Set limits on actions the tool can take.

Put it into practice: prepare cases and standards; test; record results; fix gaps; run a limited pilot only after approval. Recheck when the tool, task, or data path changes.

Keep as proof: the test plan, results, limits, correction records, and fallback steps.

Hold or pause the use when important errors remain unresolved, the fallback cannot work, or a change has not been reviewed. The aim is a dependable workflow, not a flawless-looking demo.

These proposed tests fit the spirit of NIST's voluntary AI risk guidance. NIST organizes that guidance around govern, map, measure, and manage. Those functions are not a fixed four-step sequence. NIST AI RMF.

4. Human Oversight: make review real

The goal: give the right person the facts, time, skill, and authority to review the work.

Name who checks each result and when. Put review before the result reaches a record, client, or action that needs approval. The reviewer needs access to the source facts. They must be able to edit, reject, or stop the work.

Match skills to the task. Office staff may check an appointment message. A clinician needs to review a care judgment. Billing staff should verify a billing draft against the actual service and relevant payer rules. A click on “approve” does not show that meaningful review took place.

Train staff using examples, including drafts that sound good but change a key fact. Make it easy to report repeated problems. Do not load reviewers with more work than they can check with care. Review time is part of the cost of the tool.

Put it into practice: name reviewers and backups; teach the review method; test their ability to spot meaningful errors; set a stop route and fallback.

Keep as proof: role assignments, training records, review steps, and reports that show fixes and follow-up.

Hold the use if review is required but no qualified reviewer is available, or if the reviewer cannot change the result before it has an effect.

Read The Human-in-the-Loop Principle for a full review example. NIST also calls for clear human-AI roles and responsibilities in its voluntary framework. NIST AI RMF Core, GOVERN 3.2.

5. Governance & Accountability: keep decisions clear

The goal: make sure each use has an owner, a decision trail, and ongoing review.

Governance means the rules and roles used to make decisions. A small practice may use a short shared review meeting. A larger team may use an existing quality or technology group. The point is a clear process with the right people involved.

Name a use owner who tracks the task and its results. Name who may approve, limit, pause, or end the use. Include clinical, privacy, security, operations, and billing input when those areas are affected. Seek staff and client views in ways that fit the task. A leader should own the final organizational decision, rather than leaving each user to work out the rules alone.

Keep a list of AI uses, including features added to tools already in use. Track version changes and seller notices. Set review dates and triggers for an earlier review. Changes to data use, task scope, clinical effect, or tool behavior may need fresh approval.

Plan how to leave the service. Record how data and work products can be returned or deleted under the applicable rules and contracts. Keep a way to complete needed work without the tool.

Put it into practice: maintain the use list; record decisions and conditions; assign review dates; track incidents and fixes; define the exit plan.

Keep as proof: the use record, approvals, conditions, review notes, incident log, and exit steps.

Hold the use if no one owns it, the approval scope is unclear, or the team cannot meet its conditions. NIST's voluntary framework supports clear roles, leadership responsibility, ongoing review, and plans to end system use. NIST AI RMF Core, GOVERN.

Use clear decision states

Do not turn the five areas into a single score. Strong results in one area do not cancel a missing privacy approval or absent human review. Use these proposed states for the exact task:

  • Hold: a key fact, approval, or control is missing. Name the gap, owner, and next step. Do not start live use.
  • Pilot with limits: the team has approved a small trial with a written scope, review plan, and stop rules.
  • Approved with conditions: the trial supports this use, and the required controls are in place. List users, data, tasks, limits, and review date.
  • Paused or ended: the use no longer meets its conditions or has a problem that needs action. Use the fallback and record the response.

Approval for one task does not approve the whole tool. Record separate uses when the task or data differs. No state here is a certification or proof of legal compliance.

A made-up example: drafting therapy notes

A care team wants to reduce after-hours note work. Its proposed tool drafts notes from visit information. The approved role would be drafting only. The clinician would decide and sign the final record.

For clinical fit, the team checks that the format supports its service and preserves the client's meaning. It reviews relevant evidence and records gaps. For privacy, it maps the input and logs, checks agreements and permissions, and resolves the data rules before a live pilot.

For reliability, the team tests made-up visits that include uncertain symptoms, missing facts, and changed plans. It rejects drafts that invent findings. For human oversight, clinicians practice reviewing against the source. Leaders allow time for that review. For governance, a named owner records the scope, stop rules, and review date.

The team tracks total note time, including review and fixes. It also checks errors that change meaning and staff feedback about client attention during visits. If drafts need major repair, it narrows or pauses the use. If the workflow helps and meets its conditions, it may expand in stages. This is a sample process, not evidence that a tool will save time or improve care.

Keep a decision record that staff can use

For each use, write one short record with these parts:

  1. Purpose: the work need, task, people served, and expected benefit.
  2. Scope: tool and version, users, data types, allowed actions, and limits.
  3. Five-area review: findings, proof links, open gaps, and the person handling each gap.
  4. Decision: state, approver, date, conditions, and next review date.
  5. Pilot or monitoring plan: measures, checks, feedback, and stop rules.
  6. Response plan: who handles errors, the fallback, and exit steps.

Keep evidence in approved locations. Put references to private records in the proper systems, rather than copying client details into a shared planning file. Link to the existing 25-question purchasing checklist when a new purchase is involved.

Measure benefits that matter to people

Choose measures tied to the need. For notes, track total work time and meaningful errors. For billing, track correct claims, rework, and denials with billing staff. For office work, track task time and staff effort. Include license, setup, training, review, and support costs when estimating savings.

Listen to staff and clients as well. Less typing is useful only if the full process helps care and work. Do not label time saved as better patient care without checking patient care. Do not promise higher revenue, retention, or clinical outcomes based on a demo.

Set a review date that fits the task. Recheck sooner after a serious error, a new data use, or a major tool change. The framework is meant to stay active as the work changes.

Keep this framework evolving

Version 1.0 brings the five domains together in one working guide. Future revisions should keep a dated change log that names what changed and why. Review it when laws, professional guidance, evidence, or the team's needs change. The companion governance article should use these same five domains and decision states; it should not create a second framework.

The purpose stays simple: help skilled people use useful tools well. A clear task, sound data rules, fair testing, real review, and shared ownership can give a team a better path from interest to responsible use.

Use the same five domains for new AI features

A feature change can affect several domains at once. An audio feature changes privacy needs. A source-search feature changes what must be checked for reliability. Extra training, called fine-tuning, changes the questions about training data and task fit. Review the exact feature through the five domains above. Keep the same decision process rather than treating a technical label as approval.

For safety and alignment, define what good behavior means in this use. A draft should preserve client meaning, stay within its task, protect information, and allow correction. Test those goals with varied made-up cases. Record failures and the conditions for stopping use.

AI-literacy additions reviewed October 11, 2026. See AI Terms for Care Teams for the related terms and examples.

Sources and limits

Sources checked October 5, 2026. The example and process are proposed planning tools. They do not establish a tool's performance, legal compliance, or clinical benefit. Obtain the setting-specific advice needed for live use.

Download the editable article

Back to all resources