September 17, 2026
US-China AI Agreement: What Shared Testing Would Take
Sam Altman says a US-China AI agreement could fit on one page. Here's what shared standards and testing would actually require, and what it means for builders.
Article focus
AI executives may meet Trump officials around Xi Jinping's state visit, and Sam Altman says a US-China deal on AI testing could fit on one page. Existing joint testing work suggests otherwise. Here's what such an agreement would really involve.
Section guide
A US-China AI agreement is suddenly back in the news. AI bosses are due in Washington for Xi Jinping's state visit on September 24, and Trump officials are weighing a side meeting about AI risks. Sam Altman says a deal on shared standards and testing could fit on one page. The testing work already done between friendly countries says otherwise.
Key Takeaways
- CNN reports Trump officials are considering a meeting with AI executives around Xi Jinping's state visit on September 24. Nothing is scheduled.
- Sam Altman told Fortune that Trump and Xi would deserve a Nobel Peace Prize for a deal on AI, and that a pact could fit on one page.
- Trump has called AI doom warnings a "HOAX" and rejected guardrails, which makes a safety-focused deal awkward for the White House.
- Joint testing already exists among allied countries, and its main finding so far is how hard consistent evaluation is.
- For builders, the practical effect of any agreement would arrive as standardized evaluation categories and documentation, not as new model rules.
What Is Actually on the Table?
A meeting, not a treaty. The reports describe early talks, and even the meeting isn't set.
Reported fact: CNN says Trump officials are weighing a meeting with tech and AI bosses about AI and its risks. It would sit on the edge of Xi's visit. Nothing is set yet. It's not clear if Trump would join, and Xi isn't expected to take part.
Sam Altman and Nvidia's Jensen Huang are both due at the state dinner. House Speaker Mike Johnson said a meeting with AI bosses had been in the works for a while. Any timing near Xi's visit, he said, would be chance.
A Trump official told CNN the administration is "in regular contact with AI frontier labs." It will keep working with the industry, the official said, "to cement America's AI dominance while keeping Americans safe."
The honest read: it's an odd setting. The summit is about beating China, and the side meeting is about AI risk. Still, it's the first time this White House has weighed putting lab bosses and AI risk on a China agenda.
Why Does a US-China AI Agreement Look Unlikely Right Now?
Because the US president spent the week saying the risk isn't real. That's a hard place to start a talk about safety rules.
Reported fact: In mid-September, Trump posted that "AI taking over the World, destroying Humanity, and all other things bad, is a HOAX." He called demands for guardrails part of a "SICK conspiracy." He also rejected guardrails in general. The only control needed, he said, is a strong president.
We covered the same split last week. Beijing rejected the slowdown plan, and so did the White House, for opposite reasons. Each side says holding back would hand the other a lead.
One thread runs the other way. Reports say the two governments agreed this year to set up talks between officials on AI. Talks are not a deal. But they are the channel any deal would need.
The honest read: the gap is wide between "this is a hoax" and "let's agree on testing." Lawmakers who hope the summit starts something are hoping the politics change, not just the diary.
What Would a US-China AI Agreement Have to Contain?
Far more than one page, judging by what friendly countries have found. Altman's framing sounds clean. The record makes it look hopeful.
Reported fact: Altman told Fortune that "even if just the US and China could agree on some shared standards and testing for development of this technology, I think that'd be a wonderful accomplishment." He said the terms could fit on one page.
Here's what "shared standards and testing" has required in practice between friendly governments:
| Requirement | What it involves | Status today |
|---|---|---|
| Agreed risk categories | Naming which harms get tested, such as cyber, bio, and fraud | Partly aligned among allies; separate frameworks in China |
| Common test methods | Shared prompts, scoring, and harnesses so results compare | Still being developed in joint exercises |
| Language coverage | Tests that work across languages, not just English | An open research problem |
| Access to models | Who may test which models, and before or after release | Handled by bilateral agreements with labs |
| Verification | How each side confirms the other complied | No accepted mechanism exists |
How Far Has Joint AI Testing Actually Got?
Further than most people think. Not nearly far enough to copy across a rivalry. The lesson from these tests is about how hard they are, not what they found.
Reported fact: Safety bodies from nine countries have run joint testing exercises. A recent round tested AI agents on data leaks, fraud, and cyber attacks. Pass rates ran from 33% to 57% for one model, and 14% to 35% for another. The goal was to learn how to test, not to rank models.
That range is the point. The same model scored very differently depending on how it was tested. Any deal between countries would inherit that problem.
The two countries also start from different institutions:
- United States. The old AI Safety Institute is now the Center for AI Standards and Innovation. Its framing is national security. It has signed deals with several labs to test models before release.
- China. Its standards body TC260 put out an AI Safety Governance Framework, updated in 2025. It grades risk on five levels and covers the whole life of a system. Its 2026 plans focus on grading the safety of AI apps.
- Shared ground. Both frameworks cover cyber misuse, data safety, and loss of control. The words differ. The topics overlap more than the politics suggest.
The honest read: the groundwork for narrow joint work exists on both sides. What's missing is agreement on words, on access, and on who checks what.
What Would Experts Put in a First Agreement?
Something much smaller than "shared standards." The people who study this favor small steps you can check.
Researchers in long-running US-China talks have a short list. Much of this work is run by Brookings and Tsinghua University:
- Red lines on nuclear systems. Keep AI out of nuclear command and control on both sides.
- Human control of serious cyber operations. Require a person in the loop for consequential attacks.
- An incident hotline. A named channel for AI incidents, similar to crisis lines used elsewhere.
- A terminology working group. Agree on what terms like human oversight actually mean before writing rules.
- Narrow technical risks first. Focus on cyber misuse, weapons-related misuse, and reliability failures.
The honest read: none of that caps power or slows work. It aims to avoid accidents. That may be the only part both governments want. A one-page deal is possible if the page covers a hotline and shared words, not shared testing.
What Would a US-China AI Agreement Mean for Builders?
Standard testing expectations, arriving the long way around. Governments rarely touch your internal work. But their test topics turn into the questions your customers ask.
Here's how that flows down:
- Evaluation categories become procurement questions. Enterprise buyers copy government risk categories into vendor questionnaires. Cyber misuse, data leakage, and reliability are already common.
- Documentation becomes table stakes. Model cards, system descriptions, and test records move from nice-to-have to required attachments in deals.
- Incident reporting gets formalized. If governments build reporting channels, large customers will expect you to have one too.
- Testing vocabulary converges. Shared terms make it easier to compare vendors, which means your claims get compared more directly.
- Nothing changes overnight. Even a signed agreement would take years to reach procurement. Building the habit early is cheap; retrofitting isn't.
- Model choice gets a paperwork dimension. If testing categories converge, vendors that publish results become easier to justify to a security reviewer than ones that don't.
The honest read: the useful preparation is the same work we'd recommend anyway. Keep a real evaluation set, write down what your system does and doesn't do, and log incidents. Our guides to AI agent evaluation and AI governance cover the mechanics, and EU AI Act compliance shows how formal regimes tend to land on teams.
How Do You Prepare for a US-China AI Agreement?
You don't wait for one. The work that would matter is work that pays off anyway, so treat any deal as a deadline rather than a trigger.
Four steps cover most of it:
- Write down what your system does. One page per AI feature: what it does, what data it touches, what it must never do. That's the base for any questionnaire you'll get.
- Keep an eval set from real tasks. Fifty to two hundred real examples with expected outcomes. Run it on every model change, and keep the results.
- Log what agents actually did. Store the inputs, the tool calls, and the outputs. Without logs, you can't answer questions about past behavior.
- Name an incident path. Decide who gets told, how fast, and where it's recorded. Most reporting rules start here.
The honest read: none of this depends on politics. It's the same work that makes a system debuggable, so the cost is low even if no deal ever happens. The upside is that if testing rules do arrive, you answer them with records you already keep.
There's a second reason to start now. Buyers ask these questions long before regulators do. Enterprise security reviews already ask how you test AI features and what happens when one misbehaves. Teams with answers close deals faster, which is a better argument for this work than any summit.
What Should You Watch on September 24?
Watch for mechanisms, not photos. A state dinner makes images. A deal makes named channels and agreed words.
- Whether the executives meeting happens at all. It was still unscheduled as of CNN's report.
- Any mention of a hotline or incident channel. That would be the most achievable concrete outcome.
- Language about testing or standards. Even an agreement to define terms would be a real step.
- Export control signals. Any movement on chips would tell you more about the relationship than safety language will.
- Whether AI appears in the joint statement. Its absence would confirm the topic stayed at the dinner table.
- Who shows up for China. If Beijing sends technical staff rather than only diplomats, that hints at real working-level talks.
- Any follow-up date. A scheduled next meeting matters more than warm words at this one.
The honest read: expect little. The point of watching isn't to call a breakthrough. It's to see early whether the two sides are laying the pipes any future deal would need. We'll track it next to our coverage of the AI slowdown plan and what the markets made of it.
How Van Data Team Helps Teams Stay Audit-Ready
We help teams build AI systems that can answer hard questions without a scramble. Those questions come from customers, auditors, and regulators. That means eval sets from real tasks, clear logs of what agents do, plain notes on system limits, and an incident process that leaves records.
Whatever happens at the summit, the trend is steady. More people outside your team will want proof of how your AI behaves. A US-China AI agreement would only speed that up.
Teams that already have that proof will treat each new rule as paperwork. Our work on AI agent evaluation and human review loops for production AI agents is where we'd start.
Article FAQ
Questions readers usually ask next.
These short answers clarify the practical follow-up questions that often come after the main article.
Need a similar system?
If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.
Book your free workflow review here.
Related articles
View all
Agent Reproducibility: Lessons From the ICML Challenge

Computer-Use Agents: An Engineer's Production Guide
Gemini 3.6 Flash vs Claude Opus 5: Route by Tier

