China × US Conference for AI Safety
A working-conference proposal for the practitioner tier of Trans-Pacific AI safety — artifacts, not statements.
1. Summary
Current landscape and problem statement: US–China AI safety cooperation runs through roughly five recurring venues. All of them convene the same tier of person: Turing laureates, lab directors, and standing government advisors. They are valuable and should continue. These venues, however, overlook the AI safety researchers and advocates who are not in senior or leadership roles.
Why TCAIS: it proposes a working conference for that tier, built on three commitments: it selects for career practitioners rather than principals; it produces safety-boosting artifacts and lasting connections rather than signed statements; and it operates as a collaborative program with continuous engagement rather than an annual event.
2. Value propositions
2.1 Cross-Pacific teams producing artifacts together
Trans-Pacific teams jointly build aligned technical artifacts. Participants leave each edition with committed collaborations, and the artifacts compound across editions. Judging for POC contests gives higher weights for increased cross-team and cross-nation collaboration.
Technical artifacts. Establish mutual collaborative discourse for key safety topics. Produce artifacts that document the aligned goals for:
- Shared safety benchmarks. A jointly maintained benchmark suite for agentic safety. These provide high leverage by remaining useful for participants' primary research goals regardless of the bilateral climate.
- Operationalized red lines. Technical probes that transform IDAIS-Beijing's qualitative red lines (e.g., self-replication, weapons development) into concrete, runnable capability tests using shared instrumentation.
- Cross-replication. 2–3 agentic-safety results reproduced across labs on both sides, with public documentation of all results and failures.
- Bilingual evaluation-terminology reference. A mapping of ~50 safety and alignment terms based on actual ecosystem usage, developed iteratively as engineering disputes arise during collaboration.
Project workshop track. A hands-on track where small teams build a proof-of-concept within the 2–3-day conference, focused on specific technical AI safety projects (e.g., embodied AI safety, interpretability probes, deception detection). Teams exit with a working POC, named collaborators, and a follow-up cadence for continuing the work between editions.
Motivations for participation:
- Publications. Joint artifacts and cross-replications are publishable results with novel cross-lab authorship.
- Company and institutional projects. Benchmarks and evaluations feed directly into participants' primary research, so the work stays valuable regardless of the bilateral climate.
2.2 Lower barrier of entry
Every existing Trans-Pacific venue selects for people who already hold seniority. TCAIS will facilitate meaningful conversation and collaboration for AI safety professionals. Building a strong fabric of collaboration and mutual understanding in this layer boosts systemic resilience, reduces competitive dynamics, and fosters safety culture.
Implementation: Target attendee is 0–8 years post-degree and currently working on an AI safety project — a lab safety team, an academic safety group, an eval org, or sustained public safety work. Exceptions for promising researchers transitioning into the field, contingent on funding headroom. Senior participants capped at ~10% and would attend as mentors. Selection is by affiliation, invitation, or work submissions. Full funding support is provided to boost accessibility.
2.3 Persistent, not episodic
Existing venues are annual or less, with no structured activity between sessions; relationships formed there have no channel through which to compound. TCAIS runs a continuous program with two convenings per year as checkpoints. Cohorts are additive.
- Collaborative space, not just delivery. Shared repository, low-volume bilingual technical channel, and a monthly cross-timezone research call — a standing venue for knowledge-sharing between teams, separate from artifact deadlines.
- Recommended follow-up template. A standard format for post-conference work — scope, owners, milestones, definition of done, deadlines — so teams leave with structure rather than intentions.
- Institutional buy-in as the retention mechanism. Participants have full-time jobs; follow-through depends on employers wanting the work done. Artifacts are therefore chosen to feed participants' existing research mandates, and each team leaves with a one-page brief they can take to their manager.
3. Accessibility
Friction — cost, visas, employer approval, and physical location — is the binding constraint on who can attend, and it falls asymmetrically. A convening that is nominally open but practically reachable only by well-funded participants at large labs reproduces exactly the selection problem §2.1 is meant to solve. Access design is therefore treated as a core requirement, not logistics.
- Funding. Philanthropic mainly. Candidate funders are the foundations already active here (Berggruen, Minderoo, Open Philanthropy and comparable), plus AI-safety-specific grantmakers.
- Full travel support by default. Travel and accommodation covered for all accepted participants, with opt-out for those whose employers pay. Need-blind selection: funding status is not visible to the review committee. Budget a stipend line for participants losing income to attend.
- Venue. Hosted on neutral ground. SIN or something similar.
- Visa support. Named as a program responsibility, not the participant's problem. Concretely: invitation letters issued on host-institution letterhead within 48 hours of acceptance; a 4–5 month lead time between acceptance and convening; fee reimbursement; a designated coordinator tracking each application; and a documented remote-participation fallback so a visa denial costs a participant the room, not the cohort.
- Employer approval. Scope, agenda, and output policy published in advance so internal approval at a frontier lab or a Chinese lab is a low-stakes task. No proprietary information involved, no press, no joint statements to be held against anyone.
- Language. Written outputs dual-language; working groups mixed by default. This is an access commitment rather than a headline feature, but it determines who can contribute substantively rather than just attend. Promote bilingual attendants for more effective social connections.
4. Format
| Component | Cadence | Description |
|---|---|---|
| Convening | 2/year, 2.5 days | Working sessions, not talks. ~70% of scheduled time is hands-on small-group work on committed artifacts. Day 2 includes a cross-replication block where teams swap and reproduce each other's results. |
| Async layer | Continuous | Research call, shared repository, bilingual technical channel, artifact issue tracker. |
| Cohort | 30–40, first convening | Focused on US/China collaboration; open and welcome to third-party participants. |
5. Next steps
- Institutional outreach. Contact safety organizations and conference organizers for transparency, visibility, and potential collaboration. Approach both sides in parallel — either alone will ask who the counterpart is.
- China-facing: Concordia AI, CnAISDA-affiliated groups, Shanghai AI Lab, Tsinghua AIR, WAIC safety forum organizers. Establish whether tacit non-objection from relevant authorities is sufficient or explicit approval is required.
- US- and internationally-facing: Safe AI Forum (which runs IDAIS), FAR.AI, METR, CAIS. The goal is advisory endorsement as a complementary tier, not institutional ownership.
- Funding. Draft the case against the §3 budget. Approach philanthropic funders on both sides simultaneously.
- Cohort, venue and sponsorship. Seed the participant list from people already doing safety work on both sides. Shortlist Singapore venues; confirm capacity for 30–40 and a 2.5-day working format.
- Conference arc. Agenda, workshop track structure, artifact team formation, cross-replication block, and a definition of done for each first-cycle artifact. One artifact scoped in real technical detail is the strongest single thing to show a funder.
- Logistics and safety. Venue security, medical and emergency arrangements, a named on-site point of contact, an incident policy, and a code of conduct published with the call for applications.
- Visa workstream. Confirm which host institution can issue invitation letters, map processing times in both directions, and set the acceptance deadline backwards from the longest of them.
6. Immediate action items
| # | Item | Output |
|---|---|---|
| 1 | Identify ownership. Who is working on this, at what commitment level (hours/week), and where volunteers or collaborators are needed. | Named owner and contact |
| 2 | Gather feedback and finalize the proposal. Circulate for review; resolve open design questions. | Proposal v1.0 |
| 3 | Open institutional conversations. First outreach to two or three orgs per side. | Meetings booked |
| 4 | Identify funding sources. Map grants and funders; note deadlines and lead times. | Ranked funder list with dates |
| 5 | Publish a LessWrong post. Value propositions, landscape context, project timeline. Doubles as a recruiting channel for collaborators. | Draft ready for review |