# Sovereign Clinical AI Edge Infrastructure Technical Specification for Building-Integrated Compute and Data Sovereignty for Next-Generation Medicine Specification: Draft v0.2.0 — Request for Comment (RFC) ## Abstract The delivery of modern healthcare across suburban and retail Medical Office Buildings faces an operational bottleneck: the friction between cloud-dependent AI platforms and strict regulatory, network, and performance constraints. Contemporary clinical workflows demand continuous multimodal AI — real-time ambient transcription, immediate documentation synthesis, automated triage of high-resolution imaging, and low-latency augmented-reality guidance inside a bounded end-to-end budget — yet routing Protected Health Information across public networks to hyperscale clouds introduces prohibitive latency, recurring egress penalties, an expanded Business Associate Agreement and subprocessor surface, and network saturation risk. This specification defines a new healthcare real estate asset class: a zero-egress, on-premise sovereign edge compute node built into the physical Medical Office Building, powered by tenant-owned NVIDIA L40S-class GPU clusters in acoustically isolated, liquid-cooled utility enclaves. Ambient clinical data is ingested and processed locally — never crossing the property's network boundary — delivering tenant custody of data, elimination of recurring egress cost, and a measured glass-to-photon latency budget within 13 milliseconds for interactive guidance. The specification positions zero egress as a physical and technical control within a defense-in-depth HIPAA Security Architecture pursuant to 45 CFR § 164.312, not as a substitute for the administrative, physical, and technical safeguards that rule requires. Architecture reduces the exposure surface; it does not by itself establish regulatory compliance. ## Revision History * **v0.2.0 (August 2026):** Institutional review revision. Reframes zero egress as a physical and technical control within a defense-in-depth HIPAA Security Architecture under 45 CFR § 164.312 rather than as compliance established by locality. Narrows the Business Associate Agreement claim to the elimination of third-party hyperscaler data-processor and subprocessor agreements. Restates acoustic confidentiality against ASTM E1130 Privacy Index and Articulation Index criteria in place of intelligibility assertions. Adds an explicit glass-to-photon latency budget, a Policy & Authorization Engine governing Model Context Protocol tool execution, and a persistent-artifact lifecycle acknowledging that generated clinical documentation becomes retained Protected Health Information on ingestion into the record. Recasts the valuation illustration as an underwriting sensitivity analysis, and conditions Fair Market Value arrangements on independent appraisal and commercial reasonableness under 42 CFR § 411.357. * **v0.1.0 (June 2026):** Initial Request for Comment. ## The Clinical AI Bottleneck The medical practices best positioned to leverage frontier AI — high-volume primary care, ambulatory surgery centers, outpatient imaging, orthopedics, and dermatology — are precisely the practices least able to use it as delivered. The models are capable. The regulatory, network, and performance envelope required to run them on raw, unredacted Protected Health Information inside a suburban Medical Office Building does not exist in the standard cloud delivery model. The bottleneck is not a model problem. It is an architecture problem. Contemporary clinical workflows require continuous, multimodal AI: real-time ambient acoustic transcription of the exam-room encounter, immediate clinical documentation synthesis, automated triage of high-resolution DICOM imaging, and augmented-reality guidance during minor procedures held inside a bounded glass-to-photon latency budget. Each of these is a continuous inference workload pointed directly at the most sensitive category of regulated data a practice handles. Routing that data across the public wide-area network to a hyperscale cloud provider introduces four structural failures at once. Glass-to-glass latency regularly exceeds 150 milliseconds — fatal for AR overlay and interactive guidance, where anything above roughly 15 milliseconds induces visual distortion and procedural error. Continuous 4K video, volumetric DICOM transfers, and ambient room audio trigger recurring cloud data-egress tariffs that scale with exactly the telemetry that makes the AI useful. Every outbound call carrying PHI creates third-party Business Associate Agreement liability and subprocessor exposure. And the sustained multi-gigabit uplink required makes the entire clinic dependent on an ISP circuit that, when it drops, takes clinical AI down with it. The AI-Native Medical Office Building removes this bottleneck. By establishing a localized, bare-metal GPU enclave inside the physical property — an NVIDIA L40S-class cluster in an acoustically isolated, liquid-cooled utility space — the practice executes high-throughput inference across a local network whose transport contribution is measured in single-digit milliseconds. Raw PHI never leaves the building envelope. What changes is the location of the control: instead of a procedural promise about what a vendor will do with data after it leaves the covered entity's custody, the deployment presents an inspectable physical and technical control over data that has no configured path out of the building. That control is one layer of a HIPAA Security Architecture, not a replacement for it. The Security Rule's administrative, physical, and technical safeguards continue to apply in full to hardware sited inside the practice's own walls, and a local deployment that neglects access control, audit logging, or encryption at rest is no more defensible than a cloud one. ### Cloud-Based AI vs. Sovereign Edge Compute Evaluating the transition from hyperscale cloud services to localized infrastructure requires a technical and operational comparison across network performance, compliance frameworks, financial models, and physical facility integration. | Architectural Dimension | Cloud-Based AI APIs (AWS / GCP / Azure) | Sovereign On-Premise Edge Node | Strategic & Operational Impact | | --- | --- | --- | --- | | Glass-to-Photon Latency | 150–400 ms (public WAN routing, TLS handshakes, API gateway queuing) | 13 ms budgeted end to end (3 ms capture + 1 ms local transport + 5 ms FP8 inference + 4 ms render), measured per deployment | Brings interactive AR guidance inside the ~15 ms perceptual threshold, subject to per-site verification against the budget. | | Data Egress Liabilities | Variable & cumulative ($0.09/GB egress + per-token API fees) | $0.00 / year — all telemetry contained within the physical property | Eliminates cost spikes from continuous 4K video, DICOM transfers, and room audio. | | Privacy & BAA Surface | Broad — third-party data-processor BAAs, multi-tenant isolation, undisclosed subprocessors | Narrowed — zero-egress boundary and ephemeral execution; BAA scope limited to direct local infrastructure operators | Reduces the BAA and subprocessor exposure surface; does not remove the Security Rule obligations of 45 CFR § 164.312. | | WAN Bandwidth | High overhead — dedicated multi-gigabit symmetrical circuits required | Zero WAN overhead — intra-building GbE/10GbE processes all telemetry | Clinic runs autonomously offline; no site-wide failure during ISP outages. | | MEP Requirements | Standard internet; minimal on-site hardware | 208V/415V 3-phase power, STC-55 isolation, Direct-to-Chip liquid cooling | Specialized build-out increases lease stickiness and property valuation. | | ESG & Thermal | Exhaust dissipated at remote hyperscale centers; no facility benefit | 110–130°F waste heat reclaimed into hydronic heating and snow-melt loops | Converts server heat into a building resource; lowers PUE below 1.15. | ## Who This Is For If your practice administrator has been told that ambient clinical documentation requires uploading raw exam-room audio to a third-party scribe vendor, this architecture resolves that at the infrastructure level — the audio is transcribed on a GPU in the building and never transits a network. If your compliance officer cannot sign a Business Associate Agreement broad enough to cover continuous multimodal inference on unredacted PHI, this architecture narrows the problem at the infrastructure level — no third-party hyperscaler data processor or cloud subprocessor participates in the inference chain, so the agreements still required run only to the direct local infrastructure operators the practice selects and can audit. If your radiologists wait on cloud round-trips before a preliminary AI triage score returns, this architecture resolves that at the infrastructure level — the study is triaged the moment acquisition completes, on hardware feet from the scanner. The clinical tenants this architecture is built for operate across five specialty conditions. High-volume primary care and internal medicine, where physicians spend up to two hours on EHR documentation for every hour of direct patient care. Ambulatory surgery and urgent care, where sterile-field isolation and sub-10-millisecond AR overlay decide procedural safety. Diagnostic radiology and outpatient imaging, where multi-gigabyte DICOM studies saturate WAN links and delay acute findings. Orthopedics and sports medicine, where markerless gait and kinematic analysis must fit inside a fifteen-minute appointment. And medical aesthetics and dermatology, where multi-spectral facial imaging is both the clinical asset and the most privacy-sensitive data the practice holds. The audience extends past the exam room to the people who own and finance the building. Medical Office Building owners, healthcare REITs, and commercial real estate developers converting softening suburban office stock hold the other half of this thesis: the infrastructure that makes a practice AI-native — hardened enclaves, three-phase power, liquid-cooling loops, dark fiber — is also what supports longer tenant retention, Net Operating Income expansion, and access to specialized medical-office debt pools. Whether those operating improvements translate into a lower capitalization rate depends on the transaction, the submarket, and the buyer pool at the time of sale — it is an underwriting hypothesis to be tested per asset, not a mechanical consequence of the build-out. The clinical case and the capital-markets case are the same building. The threshold for qualification is not practice size. It is a maturity condition: a practice or property owner that has moved past cloud AI experimentation and is now confronting its regulatory, latency, and egress ceiling. If the pilot worked and the production deployment stalled on a BAA review or a latency budget, this is the architecture that resolves the stall. ## How the Architecture Works ### The Sovereign Edge Pipeline The AI-Native Medical Office Building processes ambient clinical reality through an integrated hardware and protocol pipeline sited entirely inside the property. Sensor data flows from exam-room devices into local GPU memory and never crosses a public WAN boundary. Each stage is a discrete, inspectable layer of the sovereign enclave. | Pipeline Layer | Integrated Hardware & Protocol | Core Functional Mechanic | | --- | --- | --- | | 1. Ambient Spatial Sensing | Shure MXA920 ceiling array + Casambi BLE mesh (48 kHz/24-bit, AES67/Dante, PTP sync) | Dynamic beamforming isolates speaker audio; BLE Angle-of-Arrival tracks provider and patient position in real time. | | 2. Telecom Ingestion | Asterisk PBX + ARI interface (bidirectional RTP stream forking) | Ingests uncompressed audio directly into processing memory without touching persistent storage disks. | | 3. Ephemeral In-Memory Pipeline | Linux tmpfs RAM disk (/dev/shm) | Processes raw audio frames entirely in volatile RAM; buffers are released on session termination — minimizing persistent raw encounter media. Generated artifacts are persisted downstream. | | 4. Local Speech Recognition | Streaming Whisper ASR (LocalAgreement policy) on local GPU cores | Converts multi-speaker clinical dialogue into streaming text at sub-3-second latency with zero cloud egress. | | 5. Deterministic Reasoning | Memgraph C++ in-memory database (native C++, no JVM) | Traverses local patient records and medical ontologies for sub-millisecond GraphRAG contextual queries. | | 6. Sovereign Compute Enclave | NVIDIA L40S bare-metal GPU cluster (Ada Lovelace, 48 GB GDDR6, FP8 Transformer Engine) | Executes MONAI imaging triage, high-frame-rate AR guidance, and clinical LLM synthesis. | | 7. Execution & Policy Layer | Model Context Protocol (MCP) server over JSON-RPC/REST, fronted by a Policy & Authorization Engine (OPA, mTLS, fine-grained RBAC) | Exposes room acoustic profiles, spatial channels, and compute endpoints to clinical AI agents, with every tool call authorized deterministically before execution. An optional llm.txt file provides human- and agent-readable discovery only. | Featuring 48 GB of GDDR6 memory, 864 GB/s memory bandwidth, 18,176 CUDA cores, and 568 fourth-generation Tensor Cores running the FP8 Transformer Engine, a single L40S node delivers up to 1,466 TFLOPS of FP8 inference compute. Operating over a local dark-fiber or enterprise 10GbE LAN, sensor data flows directly from exam-room devices to local GPU memory without crossing a public WAN boundary. ### The Glass-to-Photon Latency Budget GPU locality is a necessary condition for interactive clinical guidance, not a sufficient one. Proximity removes wide-area transit from the path; it does not remove sensor integration time, codec and color-space conversion, scheduling jitter, or display scan-out. A specification that claims an interactive latency figure is therefore obligated to state where every millisecond is spent, and to require that the figure be measured at the display rather than inferred from the inference kernel. The budget below allocates the end-to-end path from photons entering the sensor to photons leaving the display panel. Each line is independently measurable, which makes the total falsifiable on a specific deployment instead of aspirational across all of them. $$L_{\text{total}} = L_{\text{capture}} + L_{\text{transport}} + L_{\text{inference}} + L_{\text{render}}$$ $$L_{\text{total}} = 3\,\text{ms} + 1\,\text{ms} + 5\,\text{ms} + 4\,\text{ms} = 13\,\text{ms}$$ | Budget Stage | Allocation | Dominant Contributor | Measurement Method | | --- | --- | --- | --- | | Sensor capture & exposure | 3 ms | Rolling-shutter readout, sensor integration time, MIPI serialization | Hardware timestamp at frame-ready interrupt against an external strobe reference | | Local transport & ingestion | 1 ms | 10GbE/dark-fiber transit, PTP-synchronized AES67 framing, DMA into GPU memory | PTP-correlated packet capture at both endpoints | | FP8 inference execution | 5 ms | L40S Tensor Core kernel execution, batch assembly, memory-bandwidth pressure | CUDA event instrumentation at kernel entry and exit, reported at p99 rather than mean | | Render, encode & scan-out | 4 ms | Overlay composition, display pipeline latency, panel refresh interval | High-frame-rate photodiode capture of the panel against the source strobe | | Total budgeted path | 13 ms | Sum of allocations, exclusive of scheduling-jitter headroom | End-to-end glass-to-photon measurement, verified per site | Thirteen milliseconds sits inside the approximately 15-millisecond threshold above which overlay misregistration becomes perceptible during instrument manipulation, but it does so with only two milliseconds of margin. That margin is the operative engineering constraint: a deployment that adds a display with 8 milliseconds of internal processing, batches inference requests across concurrent rooms, or permits a non-real-time kernel to preempt the inference thread will exceed the budget regardless of how close the GPU sits to the patient. Conformance requires measurement at p99 under clinical load, not a best-case figure captured on an idle node. ### The Tripartite Ownership Model The governance architecture rests on a clear separation of ownership and responsibility across three parties, each with a distinct role and none with access to what belongs to the other two. This structure is what establishes a defensible regulatory firewall between physical real estate, compute hardware custody, and clinical intelligence operations. The Landlord provisions the physical environment: the hardened subterranean or utility shell, the STC-55 acoustic isolation, the 208V/415V three-phase power envelope, the Direct-to-Chip liquid-cooling manifolds, and the dark-fiber pathways. The Landlord builds and maintains the enclave. The Landlord does not touch the tenant's compute or clinical data. The Tenant — the medical practice — owns the compute hardware outright under a Bring Your Own Silicon (BYOS) framework. Physical custody and legal title to the NVIDIA L40S cluster running inference workloads belong to the practice, installed in the practice's dedicated space, accessible only to the practice. There is no shared compute pool and no subprocessor present inside the hardware envelope. The Software Integrator deploys and operates the intelligence stack — the MCP server, the ephemeral ingestion pipeline, the Whisper and MONAI and Memgraph runtimes — binding the tenant's compute to the physical sensors and keeping the stack current. The Software Integrator operates at the software layer only. It does not hold, transmit, or access the tenant's PHI or clinical outputs. The result: no shared infrastructure anywhere in the stack, no third-party access to clinical inference, and data sovereignty that is the logical consequence of who owns what — not a policy position. ### Physical Sovereignty Deploying high-density GPU nodes in a Medical Office Building requires an overhaul of the traditional Intermediate Distribution Frame closet, which is engineered for low-density switches and small UPS units and is mechanically unsuited to the power, heat, and acoustic load of a sovereign compute enclave. Systems architects instead build a dedicated subterranean or utility node. The enclave is constructed with double-stud wall assemblies, resilient channels, and sound-dampening insulation engineered to an STC-55 acoustic isolation rating, preventing operational noise from entering adjacent clinical areas and preventing clinical conversation from leaving. Acoustic isolation is treated as a security control, not a comfort amenity, and it is specified against measurable criteria rather than asserted absolutes: STC-55 assemblies are engineered to achieve a Privacy Index greater than 95% and an Articulation Index below 0.05 under ASTM E1130 field testing for confidential speech privacy. Those figures describe the intelligibility available to a listener at the boundary under defined test conditions; they are a quantified confidentiality threshold, not a claim that audio is physically unrecoverable by an instrumented adversary. Cooling is addressed with Direct-to-Chip liquid cooling — cold plates attached directly to the GPU and CPU processors circulating a closed-loop coolant — which eliminates high-decibel chassis fans, holds stable junction temperatures, and lets high-density nodes run reliably in compact utility spaces. ### Ambient Clinical Intelligence Every clinical AI deployment built on structured inputs — typed EHR notes, post-visit summaries, dictated letters — operates on a degraded version of the encounter. Physicians spend up to two hours on EHR documentation for every hour of direct patient engagement, and the gap between what happened in the room and what got charted afterward is where clinical reasoning goes undocumented and billing codes get missed. The AI-Native Medical Office Building eliminates that gap. Exam rooms feature ceiling-mounted Shure MXA920 acoustic arrays whose steerable beams isolate speaker voices while filtering HVAC and hallway noise. Uncompressed audio routes into an ephemeral /dev/shm RAM directory on the local GPU node, where a streaming Whisper model transcribes multi-speaker dialogue at sub-3-second latency. A local Memgraph C++ GraphRAG engine links spoken terms to the clinic's local EHR records and generates structured SOAP notes with suggested ICD-10 and CPT codes before the patient leaves the suite. Because audio frames process entirely within volatile RAM and purge on session end, raw voice recordings are never written to disk or transmitted externally. This is not surveillance. It is the practice's own intelligence system, operating on the practice's own data, in the practice's own sovereign enclave, for the practice's own clinical and administrative benefit. ### Taxonomy of Medical Practices & AI-Native Workloads The sovereign edge pipeline specializes to the workload of each practice type. Primary care runs ambient transcription and GraphRAG synthesis; ambulatory surgery runs AR overlay against a published glass-to-photon budget; radiology runs local volumetric DICOM triage via MONAI; orthopedics runs markerless kinematic gait analysis; and dermatology runs multi-spectral image alignment. Each targets a distinct latency threshold and clinical outcome, and each keeps its highest-sensitivity data — voice, DICOM volumes, facial imagery — inside the building. | Practice Specialty | Primary AI Workload | Software & GPU Target | Latency Threshold | Key Clinical Outcome | | --- | --- | --- | --- | --- | | Primary Care & Internal Medicine | Ambient transcription & contextual GraphRAG synthesis | Streaming Whisper ASR + Memgraph C++ GraphRAG on L40S | < 3.0 s (streaming text) | Reduces documentation time up to 80%; auto-generates ICD-10/CPT codes. | | Ambulatory Surgery & Urgent Care | Sub-millisecond heads-up AR visual overlay & triage | Real-time computer vision + AR render engine on L40S | < 10 ms (glass-to-glass) | Preserves sterile-field isolation; eliminates AR display lag; automates operative logging. | | Diagnostic Radiology & Imaging | Local volumetric DICOM triage & screening | MONAI framework + FP8 tensor execution on bare-metal cluster | < 5.0 s (3D study triage) | Instantly flags emergent pathologies; eliminates WAN bandwidth charges. | | Orthopedics & Sports Medicine | Markerless kinematic gait analysis & motion capture | High-frame-rate multi-stream skeletal pose estimation | < 15 ms (real-time render) | Instant 3D joint kinematic reports during routine office visits. | | Medical Aesthetics & Dermatology | Multi-spectral image alignment & spatial feature mapping | Calibrated image segmentation & spatial registration models | < 1.0 s (image alignment) | Facial imagery retained solely within the tenant boundary; automated tracking of lesions and tissue volume. | ### The Four Principles > **Zero Egress.** No inference payload, ambient telemetry, or PHI crosses the property's network boundary. The inference runs on tenant-owned GPUs inside the building. The output stays there. > **The Room as the Interface.** The exam room is the primary data source. The encounter is captured at full fidelity through ceiling arrays and spatial mesh, not reconstructed from a typed note. > **The Hardened Shell.** Acoustic and physical isolation engineered to STC-55 and verified by ASTM E1130 field measurement, with Direct-to-Chip liquid cooling. The shell makes sovereignty physically inspectable rather than merely asserted in policy — and the policy layer still has to exist. > **Sovereign Compute.** Tenant-owned NVIDIA L40S silicon under a BYOS framework. No per-token billing, no third-party access, no subprocessor in the inference chain. ## Zero-Trust Physical Identity & The MCP Standard A sovereign compute environment is only as secure as its physical access logs. The AI-Native Medical Office Building fuses cryptographic door strikes with BLE spatial positioning to achieve zero-trust physical identity. When the localized acoustic array captures an execution command, the orchestration layer cross-references the speaker's spatial coordinates against the physical security ledger. If no authenticated physical entry event exists for that presence, the packet is deterministically dropped — an agent's authority is bounded by who is verifiably in the room. The physical building is abstracted into a standardized API endpoint through the Model Context Protocol (MCP). By wrapping the hardware sensor stack — uncompressed audio, spatial telemetry, physical door strikes, and local compute endpoints — into an MCP-compliant server, any authorized on-premise clinical model can query the room's physical state using native JSON-RPC tool calls, eliminating fragile custom middleware. Exposing building hardware as callable tools to a language model creates an attack surface that the transport layer does not address. A model that ingests ambient clinical dialogue is ingesting untrusted input: a phrase spoken in the room, dictated from a patient's own document, or embedded in a scanned referral can attempt to redirect the agent toward an unauthorized tool call. Prompt injection, privilege escalation through chained tool invocations, and confused-deputy execution against the door-strike or record-retrieval interfaces are the governing risks, and none of them are mitigated by keeping the model on-premises. Locality contains where data goes; it says nothing about what the agent is permitted to do. The specification therefore requires a Policy & Authorization Engine interposed between the model and the MCP server, so that no tool call reaches hardware on the model's authority alone. Every invocation is evaluated against externalized policy — Open Policy Agent or an equivalent deterministic decision engine — carrying the caller's authenticated identity, the physical presence evidence established at the door, the conformance class of the target tool, and the clinical context of the session. Transport between the agent, the policy engine, and the MCP server is mutually authenticated with mTLS and short-lived workload credentials. Authorization is fine-grained and default-deny: tools are enumerated explicitly per role, high-consequence tools require a corroborating physical-presence assertion, and the decision, its inputs, and its outcome are written to the local audit ledger before execution proceeds. Policy lives outside the model weights and outside the MCP server, which is what keeps it auditable by a reviewer and unreachable by an injected instruction. Discovery is a separate and deliberately subordinate concern. Machine-level execution is carried by MCP over JSON-RPC and by the local REST endpoints, which are the primary system-level protocols and the only paths through which state changes. An optional llm.txt file placed at the local gateway (e.g. http://edge-node.local/llm.txt) serves as a human- and agent-readable documentation convention that indexes local GPU nodes, ceiling arrays, spatial grids, and room acoustic parameters. It is a directory, not infrastructure: it confers no authority, is never treated as a trusted source of policy or capability, and a conforming deployment operates fully without it. To maintain deterministic execution during agent interactions, the local database replaces legacy JVM-bound graph databases such as Neo4j with the native C++ in-memory engine Memgraph. Operating the knowledge graph in C++ within volatile system memory eliminates Java garbage-collection pauses and enforces sub-millisecond execution boundaries for complex GraphRAG queries. Combined with ephemeral /dev/shm storage, intermediate transcription logs, video frames, and location metrics are released when the clinical session ends, minimizing the window in which raw encounter media exists in the system. Release of a volatile buffer is a strong reduction in exposure, not a cryptographic erasure proof: residual data may persist in DRAM until overwritten, and the design assumption is that the enclave's physical and access controls — not the volatility of RAM alone — protect the interval before reuse. ## Real Estate Economics & Landlord-Tenant Alignment Integrating sovereign edge compute changes the commercial real estate underwriting profile of suburban and retail Medical Office Buildings. Traditional suburban office assets face market headwinds and softening valuations, often trading at capitalization rates between 7.50% and 8.50%. Specialized Medical Office Buildings have historically transacted at tighter rates — observed ranges of roughly 6.00% to 6.50% �� owing to tenant retention, physical build-out investment, and lease stability. Those are observed market spreads across a class of assets, not a rate an individual building acquires by installing infrastructure. The value mechanism this specification claims is operational rather than mechanical. Infrastructure-enhanced build-out creates value through three channels a lender or appraiser can diligence directly: expansion of Net Operating Income through premium rents and compute capacity fees; longer effective lease duration and lower renewal risk, because a practice whose clinical workflow depends on in-building silicon faces materially higher switching costs than one whose fit-out is millwork and cabling; and access to specialized healthcare debt pools priced against licensed medical tenancy. Capitalization rates are set by the buyer pool at the moment of sale and are a function of submarket liquidity, tenant credit, lease term, and prevailing rates — none of which a landlord controls. Any cap-rate improvement is therefore an underwriting hypothesis to be sensitivity-tested per asset, and the analysis that follows is presented as a sensitivity model rather than a projection. ### Underwriting Sensitivity Analysis for Infrastructure-Enhanced MOBs The following is a sensitivity analysis, not a forecast. It illustrates how value responds to two independent variables — Net Operating Income and exit capitalization rate — and it is included so that a reader can see how much of the headline outcome depends on the rate assumption rather than on operating performance. No party to this specification represents that any particular cell will be realized. Take a 50,000-square-foot asset generating a baseline Net Operating Income of $2,000,000. Underwritten as commodity suburban office at an 8.00% cap rate, it supports a valuation near $25,000,000. Assume the sovereign edge build-out — enclave, three-phase power, liquid-cooling hookups, dark fiber — supports premium rents and compute capacity fees that lift NOI to $2,250,000. Holding the cap rate flat at 8.00%, that NOI gain alone produces roughly $28,125,000: an increase of about $3,125,000 attributable entirely to operations, and the only portion of the outcome the landlord's execution actually drives. $$V = \frac{\text{NOI}}{r_{\text{cap}}}$$ Operating contribution, holding the rate constant: $$\Delta V_{\text{NOI}} = \frac{2{,}250{,}000}{0.0800} - \frac{2{,}000{,}000}{0.0800} \approx \$3{,}125{,}000$$ | Exit Cap Rate Scenario | Valuation at $2.25M NOI | Change vs. $25M Baseline | Underwriting Interpretation | | --- | --- | --- | --- | | 8.50% — rate widening | $26,470,588 | +$1,470,588 | Operating gains partially offset by a softer exit market; the build-out still protects basis. | | 8.00% — no rate movement | $28,125,000 | +$3,125,000 | The defensible base case. Attributable solely to NOI expansion, independent of buyer sentiment. | | 7.00% — partial re-rating | $32,142,857 | +$7,142,857 | Assumes the asset is recognized as medical rather than commodity office by a competitive bidder set. | | 6.25% — full medical re-rating | $36,000,000 | +$11,000,000 | Upper bound. Requires the asset to clear at the tight end of observed MOB pricing — an outcome contingent on market conditions, not on infrastructure. | The spread across these scenarios is the point. Roughly $3.1 million of the range is earned through NOI expansion and is diligenceable from the rent roll and the compute service agreements. The remaining $7.9 million between the base case and the upper bound is a function of the exit cap rate, which the owner does not control and which no build-out guarantees. Institutional underwriting should credit the operating case, treat any re-rating as optionality rather than basis, and stress the analysis at a widened rate to confirm the investment survives an unfavorable exit. Sensitivity to the rate assumption should be disclosed to lenders and equity partners rather than compressed into a single headline valuation. The shift also improves Commercial Mortgage-Backed Securities debt underwriting. Lenders evaluate risk using the Debt Service Coverage Ratio; securing more than 50% of spatial allocation or NOI from licensed healthcare tenants unlocks institutional medical-office debt pools with lower rates (6.20%–6.50% versus 7.00%+ for standard office), longer amortization, and higher loan-to-value limits — lowering the property owner's capital cost. ### The Colocation Compute Service Model To monetize sovereign compute without violating healthcare compliance, property owners move beyond conventional square-footage leasing toward an AI-Native colocation model. | Leasing Metric | Traditional Commercial Lease | AI-Native Colocation Compute Service Model | | --- | --- | --- | | Primary Billing Metric | Dollars per rentable square foot ($/RSF/year) | Allocated compute power & infrastructure ($/kW/month) | | Capital Improvement | Tenant finances complete interior fit-out | Landlord constructs STC-55 shell, liquid loop, and power envelope | | Revenue Stability | Fixed base rent with 2.5%–3.0% annual escalations | Tiered structure combining land rent with compute capacity fees | | Tenant Relocation Risk | Moderate; tenant can relocate at lease expiration | Low; integration with localized compute and sensors binds tenant to facility | Under this model the operator leases physical real estate at market rates and bills dedicated power capacity, direct liquid-cooling hookups, high-speed local fiber, and secure space within the STC-55 enclave on a $/kW/month basis. The infrastructure integration binds the tenant to the facility far more durably than a conventional lease. ### Infrastructure, MEP & Sustainability (ESG) The energy efficiency of compute infrastructure is measured by Power Usage Effectiveness — total facility energy divided by energy delivered to compute hardware. Legacy air-cooled server closets run inefficiently, with PUE between 1.6 and 2.0. Deploying Direct-to-Chip liquid cooling within the sovereign enclave drops auxiliary cooling power dramatically, reducing facility PUE below 1.15. Liquid cooling also unlocks building-level thermal reclamation. Coolant circulating across D2C cold plates exits the rack as heated fluid between 110°F and 130°F. Rather than dissipating this energy through external cooling towers, the closed-loop system routes heated glycol through a liquid-to-liquid heat exchanger into the building's mechanical systems — feeding hydronic perimeter heating and snow-melt loops. Converting server exhaust into usable building energy reduces heating costs, elevates GRESB and ENERGY STAR ratings, and presents institutional investors with an energy-efficient healthcare asset. ## The Compliance Moat ### Architecture as a Control Layer, Not a Compliance Conclusion Compliance in the cloud is procedural. It rests on Business Associate Agreements, subprocessor audits, vendor access controls, and contractual representations about what a third party will do with PHI that has already left the covered entity's physical control. These procedures are enforceable, but they are insufficient as the sole governance mechanism for continuous AI inference on raw clinical data — because the exposure is created at the moment the data crosses the boundary, and no agreement undoes that. What the sovereign enclave contributes is a strong physical and technical control: the data has no configured path out of the building, the compute is tenant-owned, and the audit trail sits on hardware under the practice's custody. That control is durable in a way a contractual representation is not, because degrading it requires a physical or configuration act performed inside the practice's own boundary rather than a unilateral change to a vendor's policy. It is not, however, a compliance conclusion. HIPAA compliance is a program obligation assessed against the administrative, physical, and technical safeguards of the Security Rule, and no siting decision discharges it. This specification therefore positions zero egress as one layer within a defense-in-depth HIPAA Security Architecture pursuant to 45 CFR § 164.312, and requires that the remaining layers be implemented locally rather than assumed. The practical consequence is that a conforming deployment carries the same safeguard obligations a well-run cloud tenancy would, minus the third-party processor surface. - **Access control (§ 164.312(a)(1)).** Role-based access control over every inference endpoint, MCP tool, and record interface, with unique per-user identification, automatic session termination, and emergency-access procedures. Authorization decisions are externalized to the Policy & Authorization Engine and default to deny. - **Hardware root of trust.** Measured boot anchored in a TPM 2.0 or equivalent silicon root of trust, with attestation of firmware, bootloader, kernel, and GPU driver state. An enclave that cannot attest its own software stack cannot substantiate a claim about what executed on the data. - **Encryption at rest (§ 164.312(a)(2)(iv)).** AES-256 full-disk and volume-level encryption for all local persistent storage — generated artifacts, model weights, graph database state, and audit logs — with keys held in tenant-controlled hardware and escrowed under the practice's own custody, never the landlord's. - **Encryption in transit (§ 164.312(e)(1)).** Mutually authenticated TLS across every intra-enclave hop, including sensor-to-ingestion, agent-to-policy-engine, and policy-engine-to-MCP paths. Local does not mean cleartext; the LAN is inside the boundary but is not inside the trust boundary. - **Audit controls (§ 164.312(b)).** Append-only, tamper-evident logging of inference sessions, tool invocations, authorization decisions, physical access events, and artifact writes, retained for the period required by the practice's retention policy and reviewable by an examiner without vendor cooperation. - **Integrity and authentication (§ 164.312(c), (d)).** Cryptographic integrity verification of stored clinical artifacts and model weights, plus authentication of every person and workload asserting an identity to the enclave. - **Administrative safeguards (§ 164.308).** A documented risk analysis covering the enclave, workforce training on ambient capture, a sanction policy, contingency and disaster-recovery plans for tenant-owned hardware, and periodic technical evaluation. These are unaffected by locality and remain the covered entity's obligation. Stated plainly: the architecture removes a category of risk that contracts can only manage, and it does so verifiably. It does not remove the Security Rule. A deployment that treats physical custody as a substitute for access control, encryption, and audit logging has relocated its data without securing it. ### HIPAA and the Zero-Egress Boundary Under HIPAA, a covered entity remains accountable for Protected Health Information wherever it is processed. The AI-Native Medical Office Building addresses that accountability by keeping the processing inside the entity's own custody rather than by extending contractual assurances to a remote processor. Audio, video, and EHR telemetry are processed locally on tenant-owned hardware; raw health data never leaves the physical building envelope, which also removes cloud egress tariffs on high-volume image and video pipelines. The Business Associate consequence is specific and worth stating precisely, because the general version of this claim is wrong. Zero-egress architecture eliminates third-party hyperscaler data-processor Business Associate Agreements and the cloud subprocessor chains beneath them, minimizing the practice's BAA exposure surface strictly to the direct local infrastructure operators it selects. It does not eliminate Business Associate Agreements as a category. Any party that creates, receives, maintains, or transmits PHI on the practice's behalf — a managed-services provider administering the enclave, an integrator with privileged access to the orchestration layer, a landlord technician whose maintenance role touches systems processing PHI, or the vendor of a local software component with support access — remains a business associate and requires an agreement. The gain is that this set is small, locally situated, individually negotiated, and directly auditable, rather than a multi-tier subprocessor tree disclosed by reference in a vendor's public documentation. Persistence deserves the same precision. Ambient capture is processed in volatile memory and released on session end, which minimizes the persistence of raw encounter media — the largest and least useful liability in the pipeline. But the system's outputs are the point of the system, and those outputs persist: SOAP notes, finalized DICOM reports, triage scores, structured observations, and the audit records proving what occurred all become retained Protected Health Information the moment they are written to local storage or ingested into the EHR. Those artifacts are subject to the full Security Rule safeguard set enumerated above, to the practice's retention schedule, to breach-notification obligations, and to discovery. The correct claim is minimization of persistent raw media alongside deliberate, secured retention of generated clinical records — not the absence of persistent PHI. ### STARK Law & Anti-Kickback Compliance Providing shared computing infrastructure, software tools, or below-market technological benefits to medical practices introduces risk under federal healthcare law. If a landlord or affiliated health system provides high-performance GPU compute, AI software, or specialized building infrastructure to a physician practice at below-market rates, regulators may classify the discount as illegal remuneration intended to induce patient referrals under the Physician Self-Referral Law (STARK) or the Anti-Kickback Statute. Fair Market Value is a necessary condition of a defensible arrangement, not a safe harbor that compliance can rest on. Landlord-tenant compute arrangements under this specification are structured at Fair Market Value established by independent third-party appraisal, and must satisfy the further requirements that the rental-of-office-space and equipment-rental exceptions impose under 42 CFR § 411.357: a written agreement signed by the parties, a term of at least one year, a description of the premises and equipment covered, aggregate space and equipment not exceeding what is reasonable and necessary for the tenant's legitimate business purposes, and compensation set in advance that does not vary with — and is not determined in any manner that takes into account — the volume or value of referrals or other business generated between the parties. Two conditions carry particular weight in this architecture. First, commercial reasonableness: the arrangement must make sense as a business transaction on its own terms even if no referrals passed between the parties, which for a compute enclave means the capacity provisioned bears a demonstrable relationship to the tenant's clinical throughput rather than to its referral footprint. Second, strict independence from referrals: neither $/kW/month pricing, capacity allocation, tiering, escalation, nor any discount or service credit may be structured or adjusted by reference to internal or external referral volumes. Percentage-of-revenue and per-click compute pricing are excluded for this reason. The pricing model must be validated by an independent valuation firm against prevailing market rates for comparable high-density colocation and specialized medical space, with the appraisal documented contemporaneously and refreshed on a defined cycle rather than performed once at signing. The Tripartite Ownership Model supports the analysis by keeping the benefit structurally separable — the tenant owns the silicon outright and the landlord supplies the shell and base-building systems at appraised value — so that no compute benefit flows to a referring physician below market value. Where a landlord is affiliated with a health system or any potential referral source, the arrangement warrants heightened scrutiny and independent legal review; the Anti-Kickback Statute turns on intent and is not satisfied by valuation mechanics alone. Nothing in this specification is legal advice, and structures should be reviewed by qualified healthcare counsel against the parties' specific facts. ### The Institutional Firewall By establishing legal and physical separation across the three tiers — property ownership, compute hardware custody, and clinical intelligence operations — the governance model gives each safeguard a single accountable owner and writes the separation into the physical asset rather than into a service description. The firewall is what makes the safeguard set above auditable: an examiner can determine who holds which key, who can enter which room, and who can invoke which tool, without depending on a vendor's attestation. > No third-party model access to clinical inference. No hyperscaler data-processor Business Associate Agreement and no cloud subprocessor chain governing continuous multimodal PHI processing — the remaining agreements run only to the local operators the practice selects and audits directly. No pathway for patient data into a vendor's training corpus. The audit trail lives on the practice's hardware, under the practice's control, accessible only to the practice — and to the regulators it chooses to grant access. ## Engage The practices and property owners engaging with this standard now are setting the terms for how AI-native healthcare real estate gets built. The ones waiting are not holding a position — they are ceding one. ### Reference Node Visit For healthcare IT directors, practice administrators, real estate principals, and facility systems architects evaluating the AI-Native Medical Office Building standard for their own environment. Tour a reference implementation of the standard. The full stack — STC-55 enclave, tenant-owned L40S compute, ambient ceiling arrays, ephemeral ingestion, and the MCP context layer — is deployed and operational. A reference node visit is the appropriate first step for organizations evaluating the standard: a working system that can be observed, interrogated, and stress-tested against real clinical and regulatory requirements. This is not a demonstration environment. It is the production standard. ### Tenant Inquiry For medical practices and health systems requiring dedicated sovereign clinical AI infrastructure under the tripartite model. Inquire about tenancy within a qualified AI-Native Medical Office Building. Tenant deployments provide physically isolated, purpose-built sovereign compute enclaves operated under the Tripartite Ownership Model described in this specification. The Landlord provides the hardened shell, base-building systems, and software integration. The practice owns and operates the silicon. Tenancy is appropriate for practices that require dedicated, auditable, physically sovereign clinical inference without the capital and operational commitment of building and staffing an independent facility. ### Developer / RFC Contributor For clinical informaticists, systems architects, MEP engineers, and researchers engaged with the technical standard. This specification is an open technical standard under active development. Contribute technical feedback, propose amendments, or engage with the RFC process. The standard is designed to improve through deployment experience and rigorous peer review. Practices and developers operating at the frontier of regulated clinical AI generate exactly the kind of operational evidence that makes a technical standard precise and durable. If your deployment has encountered constraints or edge cases not addressed here, that input belongs in the record. ### Initialize an RFC Conversation To request a technical briefing, begin a tenant inquiry, or submit an edge-case for the RFC specification, contact the architectural principals: Location: Armonk, NY Reference Node Routing: rfc-review@ainativemedical.org Principal: Timothy Walsh Principal: Parham Alizadeh *The AI-Native Medical Office Building is a category being built. The organizations that engage now help define what it becomes.* ## Works Cited 1. The AI-Native Office — The Room as the Machine · Draft Specification (RFC), accessed June 16, 2026, https://www.ainativeoffice.org/ 2. NVIDIA L40S: Pricing, Specs, Best Uses & Where to Run (2026) — Fluence Network, accessed June 16, 2026, https://www.fluence.network/blog/nvidia-l40s/ 3. L40S GPU for AI and Graphics Performance — NVIDIA, accessed June 16, 2026, https://www.nvidia.com/en-us/data-center/l40s/ 4. The LLM.txt Directory MCP Server: Your AI's Guide to Codebases — Skywork, accessed June 16, 2026, https://skywork.ai/skypage/en/llm-directory-ai-codebases/1978668209129771008 5. What Is LLMs.txt? Guide for AI Crawlers — Similar AI, accessed June 16, 2026, https://similar.ai/guides/llms-txt/ 6. Meet llm.txt: the smart way to connect AI to the web — DataDope, accessed June 16, 2026, https://datadope.io/en/meet-llm-txt-the-smart-way-to-connect-ai-to-the-web/ 7. What is llms.txt and Why Do You Need It? — Switas Consultancy, accessed June 16, 2026, https://www.switas.com/articles/what-is-llms-txt-and-why-do-you-need-it 8. MXA920 — Ceiling Array Microphone, Product Documentation — Shure, accessed June 16, 2026, https://pubs.shure.com/view/guide/MXA920/en-US.pdf 9. How Can Asterisk Play the Real-Time Audio Stream? — Asterisk Community, accessed June 16, 2026, https://community.asterisk.org/t/how-can-asterisk-play-the-real-time-audio-stream/105166 10. Turning Whisper into Real-Time Transcription System — arXiv, accessed June 16, 2026, https://arxiv.org/html/2307.14743v2 11. Casambi System Overview, accessed June 16, 2026, https://casambi.us/wp-content/uploads/sites/2/2024/04/Casambi-System-Overview_EN_V5.0.pdf 12. GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval — arXiv, accessed June 16, 2026, https://arxiv.org/html/2605.20815v1 13. How to Build Private AI Infrastructure for Healthcare (2026 Guide) — OneSource Cloud, accessed June 24, 2026, https://www.onesourcecloud.net/blog/private-ai-infrastructure-healthcare 14. HIPAA Security Rule To Strengthen the Cybersecurity of Electronic Protected Health Information — Federal Register, accessed June 24, 2026, https://www.federalregister.gov/documents/2025/01/06/2024-30983/hipaa-security-rule-to-strengthen-the-cybersecurity-of-electronic-protected-health-information 15. Protecting Radiology Data and Devices Against Cybersecurity Threats — PMC, accessed June 24, 2026, https://pmc.ncbi.nlm.nih.gov/articles/PMC13103043/ 16. STC Rating Chart: Walls, Doors, & Windows — Commercial Acoustics, accessed June 16, 2026, https://commercial-acoustics.com/sound-advice/stc-rating-chart/ 17. Methods — GraphRAG — Microsoft Open Source, accessed June 16, 2026, https://microsoft.github.io/graphrag/index/methods/ 18. How Would Microsoft GraphRAG Work Alongside a Graph Database? — Memgraph, accessed June 16, 2026, https://memgraph.com/blog/how-microsoft-graphrag-works-with-graph-databases 19. What is Confidential Computing? Secure Data Processing — OVHcloud, accessed June 24, 2026, https://www.ovhcloud.com/en/learn/what-is-confidential-computing/ 20. Why Regulated Industries Are Mandating Sovereign AI Stacks — Oxmaint, accessed June 24, 2026, https://oxmaint.com/sap-integration/sovereign-ai-regulated-industries # Appendices ## Appendix A: The Cloud Egress Trap: The Physics and Economics of Multimodal Data Hyperscaler infrastructure is priced on an asymmetric model: inbound data transfer (ingress) is aggressively subsidized or free, while outbound data transfer (egress) is metered and billed. For most enterprise software workloads — transactional APIs, document storage, asynchronous batch processing — this pricing structure is manageable. The cost asymmetry becomes a significant architectural constraint when the workload shifts to continuous, uncompressed multimodal telemetry. The organizations that encounter this constraint are not making avoidable errors; they are running into a structural mismatch between a pricing model designed for one class of workload and an infrastructure requirement defined by a fundamentally different one. ### The Physics of Ambient Data Generation A traditional enterprise software environment relies on users consciously submitting structured data packets via keyboards or asynchronous API calls. An AI-Native Medical operates continuously, capturing ambient human interaction as raw, uncompressed data. This environment utilizes real-time spatial audio, uncompressed WebRTC video streaming, SIP telephony mapping, and continuous screen telemetry. The physics of this data generation scale exponentially and cannot be mitigated by standard compression algorithms without destroying the granular context required by advanced machine learning models. Consider the bandwidth requirements for a standard real-time communication protocol utilized in a localized collaboration space. LiveKit, an open-source WebRTC-based Selective Forwarding Unit (SFU) designed for real-time applications, demonstrates the staggering network load required to process multimodal streams.1 Benchmarking a single large video room with 150 publishers and 150 subscribers at a standard 720p resolution—even with adaptive bitrate streaming (ABR) and simulcast enabled—generates incoming throughput of 50 MBps and outgoing throughput of 93 MBps.1 When evaluating the data footprint of an ambiently recorded enterprise environment across a standard workday, the continuous flow of packets requires dedicated processing power. A single 16-core compute-optimized server managing this WebRTC traffic will experience 85% CPU utilization simply to handle the decryption, packet processing, and re-encryption required to forward these media tracks.1 The equation for daily data generation is unforgiving. A single WebRTC session utilizing standard H.264 codecs at 1280x720 resolution demands 1.25 Mbps per stream.5 If a corporate office runs twenty concurrent multimodal collaboration nodes, the data generated is measured in terabytes per day. Furthermore, processing this data via cloud architecture introduces a severe physical limitation: the latency horizon. The glass-to-glass latency in video applications, or mouth-to-ear latency in audio, represents the time required for a media packet to travel from the source device, undergo encryption, traverse the public internet, reach the cloud SFU, undergo decryption, processing, re-encryption, and travel back to the edge.2 Every geographic hop, every transit ISP network boundary, and every encryption layer adds milliseconds to the round trip. For real-time autonomous agents interacting dynamically with human speech, any latency exceeding 200 milliseconds destroys the determinism of the interaction. True AI-native architectures cannot tolerate network jitter or packet loss; the computational engine must reside adjacent to the sensor. ### The Economics of the Egress Constraint The physical latency limitations of multimodal AI are compounded by the financial architecture of public cloud egress pricing. When multimodal data is processed in the cloud, inference APIs, model weights, and continuous WebRTC streams constantly move data out of the provider's infrastructure.6 This creates a pricing structure that compounds significantly on continuously streaming, GPU-heavy workloads.6 The egress pricing schedules across major hyperscalers reflect the cost structure enterprises encounter when routing multimodal AI workloads through centralized infrastructure: | Cloud Provider | Tier Level | Internet Egress Cost per GB (USD) | Source Notes | | --- | --- | --- | --- | | AWS (EC2) | First 10 TB / Month | $0.090 | 6 | | AWS (EC2) | Next 40 TB / Month | $0.085 | 8 | | Microsoft Azure | First 10 TB / Month (Zone 1) | $0.087 | 7 | | Microsoft Azure | 10 TB - 50 TB / Month | $0.083 | 8 | | Google Cloud (GCP) | Premium Tier First 1 TB | $0.120 | 6 | | Google Cloud (GCP) | 10 TB - 50 TB / Month | $0.060 | 10 | If an enterprise office generates merely 5 terabytes of raw multimodal data daily and transmits it to an AWS-hosted inference pipeline, the return trip of processed data, augmented video, and localized knowledge graphs will aggressively trigger these egress tiers. At 150 TB of egress per month, an organization will incur over $13,000 in pure transit costs on AWS, exclusive of the actual cost of the GPU compute itself. Moving data across inter-continental boundaries via Microsoft's Premium Global Network scales up to $0.181 per GB depending on the region.9 The architectural conclusion is clear. When continuous multimodal ingestion is the baseline operational requirement, the cost-optimal path is to localize the inference engine. By deploying sovereign compute nodes on-premises, the data never traverses a public network boundary. The cloud egress cost is reduced to exactly zero. This is not a position against centralized infrastructure — it is a recognition that different workload classes have different optimal architectures, and that ambient multimodal AI inference belongs at the edge. ## Appendix B: The Space as a Sensory Organ: The Death of the Keyboard The modern enterprise is built upon a legacy ingestion bottleneck: the keyboard. Digital-native companies rely on keyboards, mice, and discrete API calls to update databases after an event has occurred. This post-hoc documentation process is fundamentally flawed and highly lossy; it strips away up to 90% of the original human context, including tonal inflection, spatial positioning, hesitation, physiological state, and collaborative overlap. The AI-Native Medical advances beyond this paradigm. Instead of forcing humans to translate their multidimensional work into flattened, structured data for a machine, the architecture transforms the physical real estate into a passive sensory organ. The physical room becomes the primary ingestion interface, capturing reality natively at the machine layer. This transition requires a complete overhaul of localized acoustic and spatial infrastructure. ### Acoustic Telemetry and Beamforming Ingestion To achieve deterministic audio capture, the physical infrastructure requires enterprise-grade networked acoustics. Consumer-grade microphones are grossly insufficient for multi-speaker, highly reverberant environments. The AI-Native Medical utilizes beamforming ceiling microphone arrays to map acoustic energy dynamically across a three-dimensional coordinate system. The Shure MXA920 ceiling array exemplifies the required standard for spatial acoustic telemetry.11 Operating via standard Power over Ethernet (PoE) and consuming a maximum of 10.1 Watts, the unit integrates directly into the enterprise local area network.11 Instead of a single omnidirectional recording that flattens audio, the MXA920 array utilizes advanced digital signal processing (DSP) to apply precise mathematical delays to multiple internal channels, electronically steering the acoustic beam in real-time to follow active talkers.14 - Acoustic Precision: The array provides up to 8 independent transmit channels and 1 automix output, capturing audio at a 48 kHz sampling rate with a 24-bit depth and a 77.5 dB dynamic range.13 - Acoustic Echo Cancellation (AEC): The hardware features up to 250 ms of AEC tail length, alongside dedicated noise reduction and automatic gain control, ensuring the raw feed is pristine before it reaches the inference layer.13 - Network Transport: This uncompressed audio is distributed across the localized network using AES67 or Dante digital audio protocols.13 Dante networking ensures strict clock synchronization via the Precision Time Protocol (PTP) and utilizes layer 3 Quality of Service (QoS) Differentiated Services Code Point (DSCP) prioritization to guarantee deterministic packet delivery.15 Because a single Dante flow can contain up to 4 audio channels, the network handles raw, uncompressed audio packets continuously, feeding them directly into local GPU nodes.15 When this raw Real-time Transport Protocol (RTP) audio stream is directed into an open-source private branch exchange (PBX) framework like Asterisk, the telephony architecture merges seamlessly with the AI architecture. Asterisk allows external media channels via its Asterisk REST Interface (ARI) to fork bidirectional real-time RTP streams directly into a localized transcription engine.16 Instead of waiting for a meeting to end, the AI-Native Medical implements a streaming variant of the Whisper ASR (Automatic Speech Recognition) model. Utilizing a LocalAgreement policy with self-adaptive latency, the Whisper-Streaming implementation achieves simultaneous, sub-3-second latency transcription on unsegmented long-form speech.17 Because the Asterisk server is local, the audio is never sent to a centralized API; it is processed directly on the localized PCIe silicon, ensuring absolute privacy and zero latency. ### Spatial Tracking and BLE Mesh Networks Audio ingestion alone is insufficient; spatial context is mandatory for true intelligence. An AI model must know not just what was said, but who said it, where they were positioned relative to visual displays, and how they moved through the environment. The AI-Native Medical tracks movement and occupancy using Bluetooth Low Energy (BLE) positioning technology deeply integrated into the architectural lighting grid. The system relies on Casambi's BLE mesh network, which acts as the spatial nervous system of the office. Casambi establishes a decentralized, self-organizing wireless mesh network where all the intelligence is replicated in every node, completely eliminating single points of failure that plague gateway-dependent systems.18 While Casambi is traditionally specified for Human Centric Lighting control, its nodes feature built-in iBeacon capabilities, broadcasting high-frequency 2.4GHz radio signals across the physical envelope.20 Traditional indoor positioning relied on Received Signal Strength Indicator (RSSI) metrics, which are highly vulnerable to multipath fading and interference, resulting in unacceptable meter-level inaccuracies.22 The AI-Native Medical discards RSSI in favor of Bluetooth 5.1 Direction Finding, specifically the Angle of Arrival (AoA) methodology.23 By deploying a constellation of multi-antenna anchors in the ceiling, the system measures the phase differences of incoming unmodulated continuous wave signals emitted by employee badges or smartphones.23 This allows the system to triangulate the precise location of any BLE tag with centimeter-level precision.24 When this raw AoA data is preprocessed and fed into localized machine learning models—such as Support Vector Machines (SVM) or K-nearest neighbors (KNN)—the spatial tracking achieves localization accuracy exceeding 96.58% in real-time environments.22 This continuous telemetry—identifying who is speaking via the Shure MXA920, where they are standing via Casambi AoA, and what digital assets are displayed on local screens—is fused into a singular, deterministic data stream. The physical room understands the temporal and spatial context of the work natively at the hardware level, rendering manual data entry entirely obsolete. ### Stateless Visual Telemetry: The Director's Cut Acoustic and spatial positioning provide the skeletal structure of collaboration, but visual telemetry provides the context. The AI-Native Medical ingests uncompressed stereoscopic video feeds (via local RTSP/ONVIF standards routed through the LiveKit SFU). Crucially, the physical hypervisor does not record video. Surveillance relies on block-storage of raw pixels. The AI-Native Medical operates a "Director's Cut" pipeline: - Uncompressed video frames stream directly into the NVIDIA L40S VRAM via GPUDirect RDMA. - The localized, native multimodal model (e.g., Inkling) parses the frame in real-time, extracting semantic reality: identifying whiteboard schematics, tracking gaze vectors, and mapping physical interactions. - The model outputs a lightweight, structured JSON graph of the event (e.g., `[Client_A] -> [Viewed] -> [Slide_4_Pricing] -> [Duration: 12s]`). - The raw video frame is deterministically overwritten in VRAM. The system perceives the physical world, translates it into a mathematical state, and permanently destroys the biometric source material in milliseconds. ## Appendix C: The Sovereign Enclave: The Architecture of the Hardened Shell Processing terabytes of uncompressed acoustic and spatial data necessitates a physical environment engineered to the standards of a military installation. The AI-Native Medical is fundamentally different from a heavily branded coworking space; it is a localized edge compute node enclosed within a mathematically verified hardened shell. The real estate itself serves as the foundational layer of the cybersecurity stack. The sovereign compute architecture relies on a strict tripartite separation of responsibilities: - The Landlord provisions the hardened architectural shell and base building infrastructure. - The Tenant owns the local PCIe inference silicon, maintaining absolute legal and physical custody of the hardware. - The Software Integrator weaves the physical sensors and digital infrastructure together, deploying the localized orchestration layer. The Software Integrator is the cross-functional implementation partnership responsible for deploying and integrating the AI-Native Medical stack — spanning physical infrastructure design, acoustic engineering, AI orchestration, and ongoing model operations. The team is assembled per deployment, drawing from specialists across infrastructure, software, real estate, and AI systems disciplines. It translates the physical sovereign enclave into a fully operational intelligence environment. ### Acoustic Sovereignty and the STC 55 Mandate Data sovereignty is instantly voided if the physical walls leak acoustic information. In a standard Class-A commercial office, demising partitions are typically constructed with 25-gauge metal studs and a single layer of 5/8-inch drywall, yielding a Sound Transmission Class (STC) rating of roughly 38 to 40.25 At this level, normal speech is easily overheard, and loud speech can be recorded by hostile actors or unauthorized devices in adjacent corridors. The AI-Native Medical specifies rigorous acoustic isolation. The baseline structural requirement for any ingestion space is STC 55. This specification aligns with the stringent criteria defined by the Intelligence Community Directive (ICD) 705 for Sensitive Compartmented Information Facilities (SCIF).26 Under ICD 705 Sound Group 4, an STC 50 perimeter is the baseline, but STC 55 is required for conference rooms and spaces where amplified audio or multiple speakers are present.27 STC 55 is a laboratory assembly rating, so the specification states its objective in field terms: assemblies engineered to achieve a Privacy Index above 95% and an Articulation Index below 0.05 under ASTM E1130 testing of the constructed room, which is the recognized criterion for confidential speech privacy. That threshold describes the intelligibility available to a listener at the boundary under defined conditions — it is not a claim that speech is rendered entirely inaudible or that the enclave constitutes an air gap against instrumented capture.25 Achieving STC 55 requires deliberate, engineered structural modifications. Adding mass is insufficient; physical decoupling is mandatory to break the structural bridge that transmits acoustic vibrations.25 | Architectural Component | Engineering Specification | Acoustic Contribution | Source Notes | | --- | --- | --- | --- | | Structural Decoupling | Staggered 2x4 studs on a 2x6 plate, or Double Stud assemblies with a 1-inch air gap. | Eliminates mechanical path for vibration. Crucial for exceeding STC 50. | 25 | | Material Damping | Constrained-Layer Drywall (viscoelastic polymer sandwiched between gypsum). | Converts acoustic vibration energy into heat. | 25 | | Cavity Absorption | Mineral wool or high-density fiberglass batts. | Breaks up standing acoustic waves within the stud bay. | 25 | | Perimeter Sealing | Continuous acoustic-grade sealant at all joints, no back-to-back electrical boxes. | Prevents flanking paths and high-frequency sound leaks. | 25 | Furthermore, the acoustic integrity of the walls is irrelevant if penetrations are compromised. A standard solid-core wood door provides a maximum of STC 35.25 The hardened shell mandates the installation of STC 50+ acoustic door assemblies. These require cam lift hinges, RF/STC fabric-over-foam perimeter seals, and adjustable silicone drop-bottoms to maintain a hermetic seal against the threshold.26 These assemblies simultaneously provide 40 dB of RF shielding against magnetic, electric, and microwave fields in the 1 KHz to 8 GHz frequency range, preventing external radio-frequency surveillance.26 ### Dedicated Infrastructure: Dark Fiber and Power Envelopes The public internet introduces variable latency and shared routing that is incompatible with deterministic enterprise intelligence requirements. The AI-Native Medical operates independently of standard commercial ISPs. It requires dedicated point-to-point dark fiber, specifically Ethernet Private Line (E-Line) architecture. This layer-2 transport protocol connects the physical office directly to localized private data repositories or failover facilities without ever traversing public routing tables or border gateway protocols (BGP). Power infrastructure must also be deliberately provisioned. Standard office IT closets are designed for low-draw networking switches. The localized edge node requires dedicated low-voltage 20-Amp power envelopes specifically engineered for high-density compute. This power must be isolated from the general HVAC and lighting grids to prevent power cycling disruptions and ensure stable thermal management for the localized silicon. ### The Compute Engine: Sovereign Silicon and the Compute Class Specification The intelligence of the AI-Native Medical relies entirely on the tenant owning and operating their own inference silicon. The architectural standard is hardware-agnostic at the system level — the appropriate silicon depends on deployment context. This specification defines two reference compute classes. #### Class 1 — PCIe Retrofit Inference (Reference: NVIDIA L40S) For retrofit deployments within existing Class-A commercial office environments, the reference compute class is PCIe-attached inference silicon operating within standard power envelopes. Large-scale centralized GPU chassis — such as 8-way HGX systems drawing 400W per GPU — require specialized liquid cooling and 480V three-phase power that standard commercial real estate cannot support.31 The NVIDIA L40S, built on the Ada Lovelace architecture, is the reference card for this class.33 As a dual-slot, full-height full-length PCIe Gen4 card drawing a maximum of 350 Watts, multiple L40S GPUs can be deployed in standard 2U or 4U rackmount servers operating within the 20-Amp, 1.5–2kW power envelopes available in most Class-A office environments.31 The L40S provides 48 GB of GDDR6 memory at 864 GB/s memory bandwidth, 18,176 CUDA cores, and 568 fourth-generation Tensor Cores.31.34 Utilizing the Transformer Engine with FP8 precision, it delivers 1,466 TFLOPS of compute.31 In practical LLM inference benchmarks, the L40S achieves 43.79 tokens per second on an 8-billion parameter model at batch size 1, and delivers more than 2x acceleration over prior architectures for RAG workloads.34.38 Because inference workloads do not require NVLink interconnects at the node level, PCIe-attached silicon is well-suited for the localized sovereign deployment. Class 1 is the appropriate specification for any retrofit environment where power and cooling infrastructure are constrained by existing base building conditions. #### Class 2 — SoC-Integrated Sovereign Compute (Reference: NVIDIA GB10 / DGX Spark) For purpose-built sovereign nodes and greenfield campus deployments, the reference compute class is SoC-integrated silicon designed specifically for dense, energy-efficient AI inference at the edge. The NVIDIA GB10 Superchip, as deployed in the DGX Spark platform, integrates Grace CPU and Blackwell GPU compute on a unified die connected via NVLink-C2C, delivering high-bandwidth, low-latency inference in a compact power envelope suited to purpose-built physical environments — without the infrastructure overhead of traditional data center GPU chassis. This class is appropriate for dedicated AI Commons node deployments, greenfield campus builds, and any deployment where the physical environment is being purpose-engineered around the compute rather than adapted to accommodate it. #### Architectural Note Both compute classes fully support the AI-Native Medical sensor stack: Dante audio ingestion via the Shure MXA920 array, Whisper-Streaming transcription via Asterisk, Casambi BLE spatial telemetry, and localized GraphRAG pipeline execution. Silicon class is determined by deployment context; the architectural specification is constant across both. This specification is a living document. Hardware capabilities in sovereign edge compute are advancing at pace. The authors will update silicon references and compute class definitions as the standard matures and deployment experience accumulates. The Software Integrator provides the software orchestration layer that binds the selected inference platform to the physical sensor array, executing the full intelligence stack independent of public cloud routing. ## Appendix D: The Intelligence Flywheel & Absolute Sovereignty: Enterprise GraphRAG The convergence of acoustic isolation, localized PCIe hardware, and ambient telemetry creates the ultimate enterprise moat: Absolute Sovereignty. Because the uncompressed data never leaves the STC 55 physical envelope and is processed directly on the tenant-owned L40S silicon, the regulatory compliance risk drops to exactly zero. Highly regulated industries—including healthcare providers managing HIPAA-protected data, quantitative hedge funds developing alpha-generating algorithms, and law firms handling privileged discovery—are currently paralyzed by the public cloud. Utilizing managed AI services from cloud hyperscalers requires aggressive data blinding, redaction, and anonymization. This preprocessing destroys the exact temporal and semantic context the AI requires to generate deep, second-order insights. Within the AI-Native Medical, organizations ingest raw, un-blinded data directly. The local node listens to a highly confidential clinical diagnostic meeting, tracks the spatial positioning of the physicians via the Casambi AoA mesh, ingests the uncompressed audio via the Shure MXA920 array, transcribes it instantly via Asterisk, and feeds the raw intelligence into a localized GraphRAG pipeline. ### Localized GraphRAG and Hybrid Knowledge Graphs Standard RAG architectures rely entirely on vector similarity search, which fetches isolated text chunks based on semantic proximity. This approach fundamentally fails when attempting to connect disparate pieces of information across massive, temporal enterprise datasets, leading to hallucinations and disconnected logic. The AI-Native Medical employs localized GraphRAG—a hybrid architectural pattern that combines the semantic understanding of vector embeddings with the deterministic, symbolic reasoning of structured knowledge graphs.39 The implementation of a localized GraphRAG pipeline, such as the methodology defined by Microsoft Research, transforms the unstructured ambient telemetry of the office into a rigorous, queryable hierarchical structure.41 This capability is transformative; it allows AI assistants to fetch specific internal reports or customer records in real-time, drastically improving trust and relevance compared to offline Business Intelligence outputs.43 The offline indexing process operates entirely on the local sovereign compute nodes, ensuring data never crosses a firewall: - Entity Extraction: The localized LLM is prompted to process the transcribed text units, extracting named entities—such as patient names, legal precedents, financial metrics, and corporate entities—and generating a precise description for each.44 - Relationship Extraction: The system parses the documents into subject-object-predicate triples (e.g., Physician X - prescribed - Medication Y), mapping the deterministic relationships between entities across all recorded text units.45 - Community Detection: The true power of GraphRAG lies in its structural organization. The knowledge graph utilizes the Leiden algorithm to detect and group entities into highly connected, meaningful clusters or "communities." This enables multi-level reasoning, allowing the AI to understand macro-trends and hierarchical summaries across the entire temporal dataset of the enterprise.42 - Vector Indexing: Finally, the communities, entities, and relationship summaries are embedded into a local vector store, enabling rapid semantic search over the entire structured graph.46 When a user or agent submits a query within the sovereign enclave, the system does not simply guess based on vector distance. It performs a local search to retrieve highly specific entity neighborhoods, and a global search that aggregates the community-level summaries, providing LLM-based answer generation that is strictly bound to the mathematical reality of the graph.42 By utilizing a native C++ in-memory graph engine (Memgraph), the tenant can execute vector similarity and multi-hop Cypher traversals in a single atomic operation without JVM Garbage Collection (GC) pauses. This deprecates JVM-bound ontology databases and keeps GraphRAG latency inside deterministic sub-millisecond boundaries on the Class 1 node. ### The Compliance Moat This architecture creates a self-reinforcing Intelligence Flywheel. Every conversation, spatial movement, and strategic meeting occurring within the hardened shell becomes structured, queryable intelligence. The temporal and medical entities are mapped perfectly without a single piece of data ever touching a public network. By maintaining the data within an air-gapped local environment, the enterprise ensures HIPAA, FDA, and SEC compliance natively at the hardware level. The intellectual property is perfectly contained. The enterprise retains absolute ownership over not just the data, but the relationships and insights generated from that data. There is no risk of model collapse, no risk of data leakage via public cloud vulnerabilities, and no reliance on third-party security protocols. ## Appendix E: The Demise of Cloud Proxies: The Imperative for Physical Sovereignty The prevailing architecture of enterprise artificial intelligence rests on a fundamentally compromised topography. The standard paradigm extracts local physical telemetry, transmits it across public routing infrastructure, and processes it within multi-tenant hyperscaler environments. This cloud-proxy model is in tension with the baseline physics of network latency, cryptographic custody, and deterministic execution. For highly regulated environments — from healthcare diagnostic facilities and defense manufacturing floors to quantitative trading desks — reliance on external API gateways introduces attack vectors and regulatory exposure that cannot be reconciled with the governing statutes.47 Application-layer governance, as currently deployed by the major cloud providers, is inherently probabilistic, bypassable, and impossible to verify at the hardware level.48 Real-time, agentic intelligence therefore requires a shift away from centralized cloud computing toward localized, bare-metal infrastructure governed by strict cryptographic boundaries. ## Appendix F: The Hypervisor for Physical Space: Architectural Topography The localized orchestration layer functions as a hypervisor for physical space. Where a traditional Type-1 hypervisor abstracts hardware resources — CPU cycles, volatile memory, block storage — for the execution of virtual machines, the orchestration layer abstracts multimodal physical telemetry — spatial audio, uncompressed stereoscopic video, and radio-frequency positioning — for autonomous agentic consumption. It is the intermediary execution layer that sits directly between the raw environmental sensors and the tenant's cryptographically isolated GPU cluster. ### E-Line Optical Topography and Network Physics To minimize latency and guarantee physical security, the telemetry transport layer rejects standard internet-facing topologies. Routing raw telemetry over ordinary IP transit introduces jitter, variable latency, and exposure to Border Gateway Protocol (BGP) hijacking. Instead, sensory data is carried over a Metro Ethernet Private Line (E-Line).49 This is a point-to-point Ethernet virtual circuit running over dedicated, physically distinct fiber-optic cable, establishing a Layer 2 architecture in which data never touches the public internet.50 The optical transport provides sub-millisecond failover and substantial bandwidth headroom, supporting port capacities from 10 Gbps up to 400 Gbps.50 Through physical network segmentation and Virtual Local Area Network (VLAN) isolation, the orchestration layer keeps the ingestion pipeline immune to external packet injection, man-in-the-middle interception, and distributed denial-of-service (DDoS) vectors. The data path runs strictly from the localized multi-sensor arrays, through the dedicated E-Line fiber, and into the isolated server vault on the premises. Compromising the data stream would require physically cutting the fiber or breaching the acoustically hardened Sovereign Shell. ### DPDK and GPUDirect RDMA: Bypassing the Kernel Network Stack At the ingestion point of the compute vault, processing raw multimodal telemetry through the standard Linux kernel network stack introduces unacceptable bottlenecks. The conventional Linux stack is interrupt-driven: when a packet arrives at the Network Interface Card (NIC), it raises a hardware interrupt, forcing the CPU to halt execution, context-switch into kernel mode, allocate an sk_buff structure, and copy the packet from kernel space to user space. At the scale of uncompressed multi-camera video and synchronous audio, this interrupt storm starves the CPU and destroys deterministic latency. To remove these bottlenecks, the orchestration layer uses the Data Plane Development Kit (DPDK) paired tightly with the gpudev library.55 DPDK Poll Mode Drivers (PMD) disable interrupt-driven networking entirely; dedicated CPU cores instead poll the ConnectX NICs for incoming packets in a continuous loop.57 The telemetry thereby bypasses the host CPU's networking stack altogether. Through GPUDirect Remote Direct Memory Access (RDMA), incoming uncompressed video frames and audio payloads are transferred directly from the NIC, over PCIe Gen4 lanes, into the contiguous GDDR6 VRAM of the NVIDIA L40S GPUs.56 GPUDirect RDMA relies on the GPU's ability to expose regions of device memory through a PCI Express Base Address Register (BAR).59 The DPDK gpudev library allocates memory pools whose payload resides strictly in GPU memory, letting the NIC transmit and receive packets using the GPU as the primary memory target.55 | Architectural Component | Traditional OS Network Stack | Localized Orchestration Layer (DPDK / GPUDirect RDMA) | | --- | --- | --- | | Packet Reception | Hardware interrupt-driven (IRQ) | Dedicated Poll Mode Driver (PMD) | | CPU Involvement | High context switching, sk_buff allocation | Zero CPU intervention in the critical data path | | Memory Destination | Host RAM → kernel space → user space → GPU | Direct to GPU VRAM via PCIe Gen4 BAR | | Latency Profile | Variable milliseconds, high jitter | Microseconds, deterministic | | Security Posture | Vulnerable to host CPU memory scraping | Cryptographically isolated within the GPU memory boundary | This GPU-centric network I/O model is an architectural necessity: it maximizes zero-packet-loss throughput at the lowest achievable latency while enforcing a hardware-based security boundary.56 Because the raw telemetry is never resident in the host CPU's memory, an entire class of side-channel memory-scraping attacks is foreclosed.60 ## Appendix G: Stateless Multimodal Routing: The Ingestion Pipeline Processing ambient reality requires an ingestion architecture that is exceptionally performant yet fundamentally stateless. The overarching mandate of the localized orchestration layer is to perceive everything and retain nothing. The system ingests raw reality, transcodes it into structured data, and then releases the source telemetry at the memory-pointer level. The orchestration layer retains zero packets. ### WebRTC Video Routing via the LiveKit SFU For visual telemetry, the orchestration layer deploys an embedded, local LiveKit Selective Forwarding Unit (SFU) directly on the bare-metal edge nodes.61 Unlike centralized cloud video APIs — which compress video to H.264, ship it over the internet, and await server-side inference — the local SFU operates on raw, low-latency feeds.61 LiveKit serves as the real-time media backbone, transporting voice and video over WebRTC.61 The SFU does not interpret, reason about, or analyze the video; its sole function is deterministic, latency-optimized routing.61 It manages session parameters over WebSockets, transports the media securely via Datagram Transport Layer Security (DTLS) and the Secure Real-time Transport Protocol (SRTP), and forwards spatial video frames to the appropriate tenant vision models.61 - Synchronous observation bundling: to satisfy the requirements of robotics and spatial-awareness policy, outgoing video frames and state packets must arrive bundled. The livekit/portal implementation appends the sender's monotonic clock timestamp (for example, timestamp_us) as packet-trailer metadata on every outgoing frame.63 This guarantees that multi-camera arrays produce perfectly synchronized observations per system tick, letting the backend vision models process aligned stereoscopic frames without jitter-induced hallucination. - Frame decoding: video streams are decoded the moment they reach the NVIDIA L40S, using the GPU's three onboard NVDEC engines.51 This bypasses CPU decoding overhead entirely. - Zero-retention mechanism: once a spatial frame has been parsed into structured contextual data — entity bounding boxes, identification hashes, coordinate mapping — by the tenant's vision model, the raw frame buffer in GPU VRAM is overwritten. No uncompressed video frame persists longer than the inference duration. ### Telephonic and Spatial Audio Forking via Asterisk PBX Acoustic telemetry — spatial microphones and telephonic inputs — is ingested through a localized Asterisk Private Branch Exchange (PBX). Traditional audio integration relies on application-layer polling such as AGI or EAGI, which operate in blocking modes with limited audio access.64 The orchestration layer replaces this with Asterisk's AudioSocket protocol and the Asterisk REST Interface (ARI) ExternalMedia channels.64 Dialplan and Stasis initiation: when an inbound audio event reaches the PBX, Asterisk answers it and routes it to a Stasis application via the dialplan (extensions.conf), handing control of the channel to the orchestration layer's ARI client.65 Snoop channel instantiation: the ARI client creates a mixing bridge and attaches a Snoop channel to passively fork the raw audio, letting the agent monitor the session bidirectionally without disrupting it.67 ExternalMedia routing: an ExternalMedia channel is instantiated; the client queries the UNICASTRTP_LOCAL_ADDRESS and UNICASTRTP_LOCAL_PORT variables to point the stream at a localized UDP port on the loopback interface (127.0.0.1).65 The channel is configured through a strict JSON payload injected via the ARI REST endpoint.69 *ARI ExternalMedia channel configuration* ``` { "channelId": "SI_EM_AUDIO_01", "app": "software_integrator", "external_host": "127.0.0.1:10000", "encapsulation": "rtp", "transport": "udp", "connection_type": "client", "format": "slin16", "direction": "both" } ``` Payload determinism: the audio format is bound to slin16 (16 kHz, 16-bit signed linear PCM).65 Converting to slin16 avoids the degradation introduced by telephony codecs such as μ-law or A-law and matches the native sample rate expected by modern speech-to-text models.69 RTP framing mechanics: the slin16 audio is framed at precise 20-millisecond intervals to prevent buffer bloat.65 At a 16,000 Hz sample rate a 20 ms frame yields exactly 320 samples; at 16-bit depth (2 bytes per sample) every RTP payload is exactly 640 bytes.65 This deterministic packet size aligns with memory-allocation limits, eliminating fragmentation and ensuring that memory boundaries are respected during DMA transfers. ### Ephemeral Ring Buffers and Streaming Whisper Processing The 640-byte audio payloads are depacketized — RTP headers stripped to isolate the raw PCM — and written into volatile tmpfs ring buffers mounted in /dev/shm (shared memory).65 This forces the operating system to allocate the buffer strictly in RAM, preventing any block-level disk I/O or swap-file caching.70 These continuous payloads stream directly into an optimized whisper.cpp instance running locally in the GPU execution space.72 Whisper processes the ambient audio in real time, using server-side Voice Activity Detection (VAD) to trigger inference boundaries and executing speech-to-text (STT) and diarization to produce structured JSON (timestamp, speaker ID, text).65 The core of the stateless mandate is enforced here: the instant the STT model yields its structured string, the /dev/shm ring-buffer pointer is advanced, releasing the raw audio payload. The raw biometric voice data is never committed to durable storage and becomes unreachable to the pipeline within milliseconds of its creation. The resulting structured JSON is handed off to the tenant's isolated data lake, where it persists as Protected Health Information under the tenant's retention controls. Stated precisely: the orchestration layer extracts the semantic reality of a room while minimizing the lifetime of the underlying raw biometric telemetry. Advancing a buffer pointer is a strong minimization control, not a cryptographic destruction proof — residual data may remain in DRAM until overwritten, and the enclave's physical and access controls protect that interval. ## Appendix H: Edge-Native Agentic Orchestration: The Orchestration Daemon Once ambient reality has been routed, transcribed, and structured into lightweight JSON by the ingestion pipeline, it requires a central logic unit to trigger autonomous action. This is the role of the orchestration daemon — a background process running continuously within the orchestration layer, acting as the deterministic bridge between spatial awareness and the tenant's Large Language Models (LLMs) and hybrid GraphRAG databases. ### Radio-Frequency Telemetry: Bluetooth Angle-of-Arrival (AoA) True spatial intelligence requires absolute coordinate mapping of physical entities within the Sovereign Shell. Audio and video supply semantic context; radio frequency supplies mathematical coordinates. The orchestration layer uses Casambi Bluetooth Angle-of-Arrival (AoA) tracking, via exposed WebSocket APIs, to generate accurate real-time spatial positioning.74 In the AoA method the tracked entity — a physical asset, an employee badge, a medical terminal — transmits a direction-finding signal from a single antenna.75 The signal carries a Link Layer field known as the Constant Tone Extension (CTE).77 The Sovereign Shell's locator devices, equipped with rapidly switched antenna arrays, receive the signal and perform In-phase and Quadrature (IQ) sampling.77 The phase difference, $\Delta\phi$, between signals arriving at two antennas separated by distance $d$ is given by the formula [76]: $$\Delta\phi = \frac{2\pi d \sin(\theta)}{\lambda}$$ Where $\lambda$ represents the signal wavelength and $\theta$ is the absolute Angle-of-Arrival. [76] By rearranging this equation, the daemon computes the precise spatial angle [76]: $$\theta = \arcsin\left(\frac{\Delta\phi \lambda}{2\pi d}\right)$$ Aggregating these angles across multiple locators within the Sovereign Shell, the daemon computes a precise 3D coordinate intersection. These coordinates stream into the daemon alongside the structured JSON transcriptions from the Whisper models, fusing semantic intent with physical location. ### Hybrid GraphRAG: Contextual Execution The orchestration daemon continuously writes this fused data — text, timestamp, coordinate space — into the tenant's hybrid Graph Retrieval-Augmented Generation (GraphRAG) architecture.80 A pure vector database is insufficient for agentic execution because it lacks ontological awareness: it can find similar text but cannot model relationships or strict hierarchical permissions. The orchestration layer therefore mandates a dual-database approach at the edge: - Qdrant (vector database): used for semantic similarity search and rapid contextual triage of transcribed text.80 To absorb high-velocity ingestion of live transcripts, Qdrant is deployed at the edge with a two-shard layout — a mutable shard for live writes and an immutable shard mapped to the HNSW (Hierarchical Navigable Small World) synced baseline.81 - Memgraph (graph database): A native C++ in-memory graph used to store complex relationships, historical state, and spatial topologies. Memgraph maps the enterprise ontology and role-based access dependencies with sub-millisecond latency. When the orchestration daemon identifies a trigger condition, it executes a hybrid retrieval. If the Qdrant database matches a spoken command — for example, "update patient file" — the daemon extracts the associated user and entity IDs and queries the Memgraph graph for the contextual relationships linked to those IDs.80 Crucially, the Memgraph graph correlates the speaker's current Casambi AoA coordinate against the authorized physical zone for clinical data access. If the user is authorized, the daemon spawns a localized agent.82 That edge-native agent retrieves the relevant graph context, processes the localized decision through the tenant's air-gapped LLM, and executes the digital API call to update the clinical-trial file.80 The orchestration is entirely deterministic. Every agentic action is constrained by physical-proximity capability ceilings and hardware-evaluated identity rules.48 If the Bluetooth AoA data places the speaker in the hallway outside the authorized acoustic perimeter, the daemon nullifies the execution request — physically preventing the action regardless of any software-level permission or API token the user may hold. Governance lives in the kernel, tied directly to physical space.48 ### Real-Time Stakeholder Augmentation The purpose of the "Director's Cut" ingestion is not historical archiving; it is real-time capability expansion. Because the orchestration daemon fuses the visual graph, the acoustic transcription, and the spatial coordinates into Memgraph natively at the edge, it can execute zero-latency reasoning loops *during* the collaboration. As external participants speak or interact with physical assets in the room, the orchestration daemon continuously queries the localized GraphRAG. If an external counterparty mentions a specific M&A precedent or hesitates on a contract clause, the localized AI instantly traverses Memgraph to find the organization's proprietary counter-arguments or related case law. These insights are pushed via encrypted WebSockets directly to the authorized internal stakeholders' localized screens or Agentic Glass interfaces in real-time. The Sovereign Shell does not just protect the organization's intelligence; it actively weaponizes that intelligence, feeding the internal team the exact proprietary context they need at the exact millisecond the negotiation requires it. ## Appendix I: Cryptographic Isolation and the Zero-Trust Moat In highly regulated domains, data governance is not a matter of corporate preference; it is a matter of federal statute and civil liability. Deploying omnipresent sensory AI in these settings demands mathematical verifiability that data cannot be extracted, compromised, or retained outside defined regulatory bounds. The localized orchestration layer's stateless architecture is the verifiable mechanism by which HIPAA, FDA, and SEC mandates can be satisfied simultaneously without constraining the system's autonomous capability. ### Bring Your Own Silicon (BYOS) Security Model The boundary between Software Integrator orchestration and tenant data custody is absolute. The Software Integrator enforces a strict "Bring Your Own Silicon" (BYOS) model: the localized orchestration layer provides the stateless routing, parsing, and execution logic, while the tenant retains physical ownership of the hardware, the cryptographic keys, and the resulting structured data lakes. The computational engine of this architecture is the NVIDIA L40S GPU.51 Chosen for its independence from forced hyperscaler interconnects and its versatility in edge deployment, the L40S balances inference, graphics, and video processing.51 Built on the Ada Lovelace architecture, it provides 48 GB of GDDR6 memory, 18,176 CUDA cores, and 568 fourth-generation Tensor Cores.51 Security in this environment rests on silicon physics rather than operating-system policy. The L40S is Network Equipment-Building System (NEBS) Level 3 ready and features Secure Boot with a hardware Root of Trust.51 - Secure Boot: prevents unauthorized firmware modification, guaranteeing that the power-on execution environment matches the verified cryptographic hash.83 - Confidential computing: the architecture leverages confidential-computing paradigms to protect data in use.83 Hardware-based isolation and encryption ensure that applications, LLMs, and Whisper models are processed within Trusted Execution Environments (TEEs), or enclaves.84 Even if the host OS is compromised by an advanced persistent threat, telemetry resident in GPU VRAM remains cryptographically sealed and inaccessible.83 Under the BYOS model the Software Integrator initiates the Trusted Execution Environment and routes the telemetry, but the enclave is sealed with keys managed entirely by the tenant. The Software Integrator operates the pipes; the tenant holds the cryptographic lock to the processing chamber. ### Compliance Mapping: Healthcare, Defense, and Quantitative Funds The architectural constraints of this approach map directly onto the compliance requirements of the most heavily regulated industries. | Industry Domain | Core Regulatory Mandate | Architectural Solution | | --- | --- | --- | | Healthcare | HIPAA (45 CFR Part 164) — transmission security, ePHI safeguards | Stateless tmpfs audio destruction, E-Line fiber transit, L40S TEE enclaves | | Pharma, Defense | FDA (21 CFR Part 11) — non-repudiation, timestamped audit trails | GraphRAG localized state logging, deterministic AoA tracking, isolated LLM execution | | Finance, Trading | SEC (Rule 17a-4) — immutable WORM storage, communication logs | Hardware-enforced zero-cloud exfiltration, local immutable structured logs via the daemon | Healthcare — HIPAA and 45 CFR Part 164: under the HIPAA Security Rule (45 CFR Part 164), covered entities must implement rigid technical safeguards — access controls, integrity controls, and transmission security for all electronic Protected Health Information (ePHI).60 Cloud deployments introduce unacceptable multi-tenant risk: shared GPU memory across cloud instances is exposed to side-channel attack, and memory states are rarely wiped between hyperscaler jobs.60 The Software Integrator enforces compliance through hardware isolation of the L40S nodes.83 Strict E-Line segmentation, combined with /dev/shm tmpfs ring buffers that deterministically destroy raw voice telemetry milliseconds after ingestion, ensures biometric data never becomes ePHI at rest.60 The localized orchestration layer operates as a true air gap, satisfying the technical-safeguard mandates of 45 CFR § 164.312 without elaborate cloud Business Associate Agreement (BAA) webs.87 Defense and pharmaceutical manufacturing — FDA 21 CFR Part 11: for biotechnology and defense manufacturing, 21 CFR Part 11 requires secure, computer-generated, timestamped audit trails for all actions on electronic records and signatures.88 Any AI system executing quality control or predictive maintenance must keep its decisions traceable, auditable, and unalterable.47 Sending batch records or ITAR-restricted assembly telemetry to a hyperscaler violates those integrity constraints because the data crosses boundaries outside the manufacturer's control.47 The BYOS approach lets the tenant run validated, locked models directly on the factory floor.47 The orchestration daemon routes system logs and agentic execution graphs into the local Memgraph database.80 The result is a cryptographically signed graph of exactly who requested an action, where they stood (via RF AoA data), what the model parsed, and when it executed — fulfilling the audit-trail mandate of 21 CFR Part 11, subsection 10(e), natively within the edge infrastructure.88 Quantitative finance — SEC Rule 17a-4: for broker-dealers and quantitative trading firms, SEC Rule 17a-4 requires that all business communications be retained complete, accurate, and unalterable.91 The rule mandates either Write Once, Read Many (WORM) storage or an audit-trail system that logs every modification, preventing destruction of evidence related to market manipulation or insider trading.91 Extracting voice telemetry from a trading floor to a cloud transcription API risks severe non-compliance, particularly around "off-channel" communications.93 The Software Integrator ingests trading-floor audio locally through Asterisk, parses it with the isolated Whisper model, and writes the structured text directly to the firm's localized WORM array. The Software Integrator touches the packets for routing but holds no key to write, alter, or delete the destination database; the firm retains absolute custody and a provable, continuous audit trail of all floor intelligence without exposing a single proprietary algorithm or conversation to the open internet.47 ## Appendix J: System Mandate: Bare-Metal PCIe Node Deployment Protocol Deploying the Software Integrator is less an installation than a fusing of silicon and telemetry. The software executes directly above the bare-metal Linux kernel and requires uncompromising control over PCIe lanes, IOMMU groups, and CPU-core isolation to guarantee deterministic, sub-millisecond execution. To deploy the localized orchestration layer onto a tenant node equipped with NVIDIA L40S PCIe accelerators, the following sequence is executed precisely. ### I. GRUB Kernel Parameter Configuration The host operating system is partitioned at the kernel boot level to reserve dedicated resources for the orchestration-layer components and to isolate the GPU hardware for Data Plane Development Kit (DPDK) and Virtual Function I/O (VFIO) mapping. The /etc/default/grub configuration appends the following parameters to the GRUB_CMDLINE_LINUX_DEFAULT string.95 */etc/default/grub — GRUB_CMDLINE_LINUX_DEFAULT* ``` # GRUB configuration requirements intel_iommu=on iommu=pt pci=realloc noats vfio-pci.ids=10de:26f5,10de:22ba isolcpus=2-15 ``` - IOMMU activation (intel_iommu=on iommu=pt): hardware-assisted I/O memory management is enabled and set to passthrough (pt), letting PCIe devices bypass host-OS DMA translation and granting the orchestration layer the direct memory access required for zero-copy telemetry transfer from the ConnectX NIC to the L40S. - PCIe resource reallocation (pci=realloc): forces the kernel to reallocate PCI bridge resources, which is required to accommodate the 48 GB BAR memory window of the NVIDIA L40S and to ensure contiguous allocation for GPUDirect RDMA. If the BIOS allocation is too small for the child devices, the kernel resizes the BAR dynamically.96 - Address Translation Services disablement (noats): disables PCIe ATS (Address Translation Services) and the IOMMU device IOTLB.97 ATS introduces variable latency in translation lookaside buffers; for deterministic edge processing of live audio and video, memory translation must be statically pinned. - Hardware binding to VFIO (vfio-pci.ids=10de:26f5,10de:22ba): example device IDs for the L40S GPU and its associated HD-audio endpoint.95 This unbinds the NVIDIA GPUs from the default nouveau or proprietary driver during boot, capturing the devices with the vfio-pci stub driver.95 The orchestration layer then asserts control over them from userspace via DPDK. - CPU-core isolation (isolcpus=2-15): removes the specified logical cores from the kernel's Symmetric Multiprocessing (SMP) balancing and scheduler.96 These cores are dedicated to the LiveKit SFU routing threads, the Asterisk ExternalMedia event loops, and the DPDK polling drivers, guaranteeing zero context-switching interruptions during telemetry ingestion. ### II. Execution Environment Initialization After the kernel parameters are configured and grub-mkconfig regenerates the bootloader, the system reboots and initializes the localized orchestration-layer runtime.95 *Stateless tmpfs mount* ``` # tmpfs mount for stateless execution mount -t tmpfs -o size=1G,mode=1777 tmpfs /dev/shm ``` - Memory provisioning: the volatile tmpfs file system is mounted strictly for audio-pipeline ingestion, satisfying the stateless-processing mandate. This provides the 1 GB ephemeral ring buffer required by the whisper.cpp inference engine and guarantees that no audio data is ever written to non-volatile block storage.70 - DPDK binding: using the dpdk-devbind.py utility, the local ConnectX network interfaces are bound to the vfio-pci driver, detaching the NICs from the Linux kernel TCP/IP stack so the PMD can assume control. - Daemon invocation: the orchestration daemon is initialized within the Trusted Execution Environment. It establishes the local WebSocket listener for the Asterisk PBX, initializes the LiveKit SFU for WebRTC traffic, and mounts the Memgraph and Qdrant GraphRAG connections.64 *Once the initialization sequence completes, the node transitions into a fully air-gapped, stateless orchestration state. The ambient reality of the physical room is mapped directly onto localized silicon, governed by cryptographic isolation and operating without dependency on external cloud architecture. The gain is structural rather than incremental: when inference sits adjacent to the sensor, latency, custody, and compliance resolve together rather than in tension.* ### III. In-Memory Graph Ontology Execution (Memgraph) The architectural standard mandates Memgraph, a native C++ in-memory graph engine, deployed directly on the bare-metal Class 1 node. ### Memgraph Configuration Parameters The daemon configuration explicitly disables external network binding and defines hard memory ceilings to protect GDDR6 VRAM transfers. */etc/memgraph/memgraph.conf* ``` --host=127.0.0.1 --port=7687 --telemetry-enabled=false --send-telemetry=false --memory-limit=65536 --memory-warning-threshold=80 --data-directory=/var/lib/memgraph ``` ### Systemd Daemon and NUMA Topology To achieve sub-millisecond GraphRAG traversal without CPU cache misses, Memgraph must not span Non-Uniform Memory Access (NUMA) nodes. */etc/systemd/system/memgraph.service* ``` [Unit] Description=Memgraph In-Memory Edge Ontology After=network.target [Service] Type=simple User=memgraph ExecStart=/usr/bin/numactl --cpunodebind=0 --membind=0 --physcpubind=16-23 /usr/lib/memgraph/memgraph --config=/etc/memgraph/memgraph.conf LimitNOFILE=65535 LimitMEMLOCK=infinity OOMScoreAdjust=-500 NoNewPrivileges=yes ProtectSystem=full PrivateTmp=yes RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6 Restart=on-failure RestartSec=5s [Install] WantedBy=multi-user.target ``` ## Appendix K: Open-Weights Architectures and Sub-Dimensional Memory Compression The architectural viability of the AI-Native Medical requires that the localized silicon not only ingests ambient telemetry but reasons over it at parity with frontier hyperscaler models. Historically, this presented a memory-bound limitation. Two distinct breakthroughs—Sparse Mixture-of-Experts (MoE) architectures and data-oblivious vector quantization—have permanently collapsed this constraint. ### The Native Multimodal Imperative: Sparse MoE Execution The standard relies on open-weights, native multimodal models engineered on a Sparse Mixture-of-Experts (MoE) architecture (e.g., Inkling). A sparse MoE model selectively activates only a highly specialized subset of its neural network per inference. This allows massive models (975B parameters) to run efficiently with only 41B active parameters, perfectly mapping to the GDDR6 VRAM boundaries of the Class 1 Compute Specification (NVIDIA L40S) without exceeding the 20-Amp thermodynamic threshold. ### Storage Abstraction: Vector Data Obliviousness Absolute sovereignty dictates that the enterprise knowledge graph must reside entirely on local silicon. The orchestration layer employs data-oblivious quantization algorithms (implemented via libraries like `turbovec`). By compressing dense `float32` vector embeddings down to 2-4 bits per dimension, the system achieves a 16x compression ratio with near-zero degradation in retrieval accuracy. This frees up critical GPU memory, allowing the tenant's localized MoE model weights and their entire compressed vector history to co-reside safely within the same Trusted Execution Environment (TEE). --- # Normative Requirements & Conformance > The normative requirements of the AI-Native Medical specification, stated as numbered RFC 2119 clauses, with three conformance classes an implementation may claim against. The key words MUST, MUST NOT, SHOULD, SHOULD NOT, and MAY are to be interpreted as described in IETF RFC 2119. - MUST: An absolute requirement. A deployment that does not satisfy a MUST clause applicable to its claimed conformance class does not conform to this specification. - MUST NOT: An absolute prohibition. The named capability, path, or practice is required to be absent — not merely disabled by configuration or forbidden by policy. - SHOULD: A strong recommendation. Valid reasons may exist to deviate in particular circumstances, but the full implications must be understood and the deviation documented. - SHOULD NOT: A strong discouragement. The behavior is permitted only where its consequences have been examined and accepted in writing. - MAY: Truly optional. An implementation that omits a MAY clause remains fully conformant, and one that includes it must interoperate with one that does not. ## Conformance classes ### Class A — Sovereign Ambient Enclave The complete architecture. Tenant-owned inference hardware inside an acoustically engineered enclave, with continuous ambient ingestion, physical identity enforcement, and no egress path for inference payloads. Applicability: Claimed by deployments serving regulated practice areas where the spoken record is the primary asset: transaction teams, litigation groups, clinical review, and investment committees. Excludes: A Class A claim requires every clause in this specification marked applicable to Class A, including the full acoustic and ambient-ingestion requirements. Partial ambient capability is a Class B deployment, not a reduced Class A. ### Class B — Sovereign Compute Enclave Tenant-owned inference hardware inside a declared demarcation boundary with no egress path, but without continuous ambient sensory ingestion. Input arrives through conventional interfaces. Applicability: Claimed by organizations that require sovereign inference and zero egress but are not prepared to operate continuous ambient capture, whether for works-council, jurisdictional, or cultural reasons. Excludes: A Class B deployment makes no ambient-intelligence claim and must not be described as capturing the spoken record. Acoustic clauses apply only insofar as they protect displayed and audible material. ### Class C — AI-Ready Shell Building infrastructure prepared to host a Class A or Class B enclave — power, cooling, structural, and pathway capacity verified — with no tenant compute installed and no inference occurring. Applicability: Claimed by property owners and developers documenting readiness in advance of a tenant. Class C is a property-layer claim about capability, not an operational claim about workloads. Excludes: A Class C claim conveys nothing about data handling, because no data is processed. Class C must never be represented as sovereign inference, and a Class C shell must not be marketed as an AI-Native Medical in operation. ## Requirements ### 1. Conformance & Terminology How a claim of conformance is made, scoped, and withdrawn. These clauses govern the use of the specification itself rather than the architecture it describes. #### ANM-1.1 [MUST] [Class A, B, C] An implementation claiming conformance to the AI-Native Medical specification MUST declare exactly one conformance class — A, B, or C — together with the specification version against which the claim is made. Rationale: An undifferentiated claim of conformance is unfalsifiable. Naming a class and a version makes the claim reviewable and lets it expire honestly as the specification advances. Verification: The claim is published in writing and names both the class and the version. Permalink: https://www.ainativemedical.org/conformance/#anm-1.1 #### ANM-1.2 [MUST] [Class A, B, C] A conformance claim MUST identify the specific physical premises to which it applies, and MUST NOT be stated at the level of an organization, a product, or a portfolio. Rationale: Conformance in this specification is a property of a room and the hardware inside it. An organization-wide claim would assert something the architecture cannot guarantee across sites. Permalink: https://www.ainativemedical.org/conformance/#anm-1.2 #### ANM-1.3 [MUST NOT] [Class A, B, C] An implementation MUST NOT describe a conformance claim as certification, accreditation, or independent review unless an independent conformance body has been established and has issued that finding. Rationale: No such body exists at the time of this revision. All current claims are self-declared, and representing them otherwise would misstate their weight. Permalink: https://www.ainativemedical.org/conformance/#anm-1.3 #### ANM-1.4 [MUST] [Class A, B, C] The keywords MUST, MUST NOT, SHOULD, SHOULD NOT, and MAY in the AI-Native Medical specification MUST be interpreted as described in RFC 2119. Permalink: https://www.ainativemedical.org/conformance/#anm-1.4 #### ANM-1.5 [MUST] [Class A, B, C] A deployment that ceases to satisfy any MUST clause applicable to its declared class MUST withdraw or downgrade its conformance claim before continuing to represent itself as conformant. Rationale: Conformance describes an operating condition, not a milestone once achieved. Hardware is removed, boundaries are redrawn, and claims must track those changes. Permalink: https://www.ainativemedical.org/conformance/#anm-1.5 ### 2. The Demarcation Boundary Every other requirement in this specification is evaluated against a boundary. These clauses require that the boundary be declared explicitly, enumerated exhaustively, and kept inspectable. Formalizes: Cryptographic Isolation and the Zero-Trust Moat (https://www.ainativemedical.org/sections/isolation/) #### ANM-2.1 [MUST] [Class A, B] A conforming deployment MUST declare a demarcation boundary that is simultaneously physical and logical, identifying the rooms, racks, and network segments inside which tenant data is processed. Rationale: A boundary that exists only as a network diagram cannot support a claim grounded in physical custody. The declaration must be walkable. Verification: A written boundary declaration exists and can be reconciled against a site plan. Permalink: https://www.ainativemedical.org/conformance/#anm-2.1 #### ANM-2.2 [MUST] [Class A, B] A conforming deployment MUST maintain a current and exhaustive enumeration of every network path that crosses its demarcation boundary, including management, telemetry, licensing, update, and out-of-band paths. Rationale: Egress claims fail at the paths nobody counted. Management and telemetry channels are the usual omissions, and both are capable of carrying payload. Verification: The enumeration is complete against an independent scan of the boundary and is dated within the current review period. Permalink: https://www.ainativemedical.org/conformance/#anm-2.2 #### ANM-2.3 [MUST] [Class A, B] Each boundary-crossing path enumerated under ANM-2.2 MUST be annotated with the categories of data it is capable of carrying, and MUST be justified against the deployment's operating requirements. Permalink: https://www.ainativemedical.org/conformance/#anm-2.3 #### ANM-2.4 [MUST NOT] [Class A, B] A conforming deployment MUST NOT rely on a boundary whose enforcement depends on a control plane operated outside that boundary. Rationale: A boundary policed from outside is a boundary held at another party's discretion, which is the dependency this specification exists to remove. Permalink: https://www.ainativemedical.org/conformance/#anm-2.4 #### ANM-2.5 [SHOULD] [Class A, B] A conforming deployment SHOULD be able to continue serving inference for a defined minimum interval with all boundary-crossing paths severed, and SHOULD document that interval. Rationale: Survivability under full disconnection is the practical test of sovereignty. A deployment that halts when the uplink drops was never independent of it. Permalink: https://www.ainativemedical.org/conformance/#anm-2.5 #### ANM-2.6 [MUST] [Class C] A Class C shell MUST declare the boundary a future enclave is intended to occupy, and MUST state plainly that no demarcation boundary is presently in force because no tenant compute is installed. Permalink: https://www.ainativemedical.org/conformance/#anm-2.6 ### 3. Data Movement & Egress The zero-egress property, stated as a prohibition on paths rather than a preference for behavior. A control that could be reconfigured to permit egress does not satisfy this chapter. Formalizes: The Cloud Egress Trap: The Physics and Economics of Multimodal Data (https://www.ainativemedical.org/sections/egress/) #### ANM-3.1 [MUST NOT] [Class A, B] A conforming deployment MUST NOT transmit inference payloads — prompts, retrieved context, intermediate representations, embeddings, or generated outputs — across its demarcation boundary during normal operation. Rationale: This is the specification's central prohibition. Embeddings and intermediate representations are named explicitly because they are frequently treated as non-sensitive despite being derived directly from privileged material. Verification: Egress monitoring over a representative operating period shows no payload-bearing flow across any enumerated path. Permalink: https://www.ainativemedical.org/conformance/#anm-3.1 #### ANM-3.2 [MUST] [Class A, B] The absence of an egress path for inference payloads MUST be a structural property of a conforming deployment rather than a policy, feature flag, or configuration setting that a privileged operator could reverse. Rationale: A prohibition that can be lifted by changing a setting is a procedural control wearing structural language, and it collapses under the examination this architecture is meant to withstand. Permalink: https://www.ainativemedical.org/conformance/#anm-3.2 #### ANM-3.3 [MUST NOT] [Class A, B] A conforming deployment MUST NOT send tenant-derived telemetry, usage analytics, error payloads, or diagnostic samples to any party outside its demarcation boundary. Rationale: Diagnostic exhaust is the most common unexamined egress channel, and stack traces and error payloads routinely contain the exact material the boundary exists to hold. Permalink: https://www.ainativemedical.org/conformance/#anm-3.3 #### ANM-3.4 [MAY] [Class A, B] A conforming deployment MAY transmit aggregate operational counters that contain no tenant-derived content, provided each such counter is enumerated under ANM-2.2 and disclosed to the tenant. Rationale: Sovereignty need not preclude knowing whether a fan is failing. The requirement is that the exception be named rather than assumed. Permalink: https://www.ainativemedical.org/conformance/#anm-3.4 #### ANM-3.5 [MUST] [Class A, B] Model weights, container images, and software updates entering a conforming deployment MUST be verified against a cryptographic signature before installation, and the verification MUST be performed inside the demarcation boundary. Rationale: Inbound supply chain is the boundary's remaining exposure once egress is closed. Verifying outside the boundary reintroduces the trust dependency. Permalink: https://www.ainativemedical.org/conformance/#anm-3.5 #### ANM-3.6 [SHOULD NOT] [Class A, B] A conforming deployment SHOULD NOT depend on an external service for any function on the critical path of inference, including authentication, license validation, model retrieval, or rate authorization. Permalink: https://www.ainativemedical.org/conformance/#anm-3.6 #### ANM-3.7 [MUST] [Class A, B] A conforming deployment MUST implement the administrative, physical, and technical safeguards required by 45 CFR §§ 164.308 and 164.312 within the demarcation boundary, and MUST NOT rely on the absence of an egress path as a substitute for them. At minimum this includes unique user identification with role-based access control, encryption of Protected Health Information at rest, mutually authenticated encryption in transit across intra-enclave hops, append-only audit logging, integrity verification of stored artifacts and model weights, a hardware root of trust with measured boot, and a documented risk analysis naming the enclave in scope. Rationale: Zero egress is one physical and technical control inside a defense-in-depth architecture. The Security Rule applies in full to hardware sited within the practice's own walls, and a local deployment that neglects access control, encryption, or audit logging is no more defensible than a cloud one. Verification: Evidence for each enumerated safeguard, together with a current risk analysis whose scope statement names the enclave. Permalink: https://www.ainativemedical.org/conformance/#anm-3.7 #### ANM-3.8 [MUST NOT] [Class A, B, C] An implementation MUST NOT represent zero-egress architecture, on-premises siting, or tenant hardware ownership as establishing HIPAA compliance, and MUST NOT describe compliance as an architectural property of the real estate. Rationale: Compliance is a program obligation assessed against a covered entity's safeguards, not a conclusion derivable from network topology. Presenting architecture as compliance invites a reviewer to substitute a siting decision for a risk analysis. Permalink: https://www.ainativemedical.org/conformance/#anm-3.8 #### ANM-3.9 [MUST] [Class A, B] An implementation claiming reduced Business Associate Agreement exposure MUST scope that claim to the elimination of third-party hyperscaler data processors and their subprocessor chains, and MUST maintain Business Associate Agreements with every party that creates, receives, maintains, or transmits Protected Health Information on the covered entity's behalf, including managed-service providers, integrators holding privileged access, local software vendors with support access, and property personnel whose maintenance role reaches systems processing Protected Health Information. Rationale: Locality removes a category of business associate; it does not remove the category. The defensible claim is a small, locally situated, individually auditable set of agreements rather than their absence. Verification: A current business-associate register with executed agreements, reconciled against every party holding privileged access to the enclave. Permalink: https://www.ainativemedical.org/conformance/#anm-3.9 #### ANM-3.10 [MUST] [Class A, B] A conforming deployment MUST distinguish minimization of raw ambient capture from retention of generated artifacts, and MUST treat clinical documentation, finalized reports, triage outputs, structured observations, and audit records as persistent Protected Health Information subject to ANM-3.7, the tenant's retention schedule, and applicable breach-notification obligations. An implementation MUST NOT claim that no persistent Protected Health Information exists. Rationale: The system's outputs are the purpose of the system, and they persist. Releasing a volatile capture buffer minimizes raw media exposure; it is not a cryptographic erasure proof, and it says nothing about the records written downstream. Verification: A documented artifact lifecycle identifying each persisted class, its storage location inside the boundary, its encryption state, and its retention period. Permalink: https://www.ainativemedical.org/conformance/#anm-3.10 ### 4. Compute & Siting Where inference executes, on whose hardware, and under what failure and dependency conditions. These clauses establish that sovereignty is a property of physical custody. Formalizes: How the Architecture Works (https://www.ainativemedical.org/sections/architecture/) #### ANM-4.1 [MUST] [Class A, B] Inference in a conforming deployment MUST execute on accelerator hardware physically located inside the declared demarcation boundary. Verification: Hardware inventory reconciles to the boundary declaration and to physical inspection. Permalink: https://www.ainativemedical.org/conformance/#anm-4.1 #### ANM-4.2 [MUST] [Class A, B] The tenant MUST hold outright ownership of the accelerator hardware, the storage media, the inference data, and all model outputs produced within a conforming deployment. Rationale: Ownership rather than lease or license is what makes the tenant's custody claim survive the insolvency, acquisition, or policy change of any counterparty. Permalink: https://www.ainativemedical.org/conformance/#anm-4.2 #### ANM-4.3 [MUST NOT] [Class A, B] A conforming deployment MUST NOT route any portion of an inference request to a model endpoint hosted outside its demarcation boundary, including for overflow capacity, fallback, quality comparison, or evaluation. Rationale: Hybrid routing defeats the entire architecture while preserving its vocabulary. A single fallback path to a hosted endpoint makes every prior guarantee conditional. Permalink: https://www.ainativemedical.org/conformance/#anm-4.3 #### ANM-4.4 [MUST] [Class A, B, C] A conforming deployment MUST be provisioned with power and thermal capacity sufficient to sustain its accelerator hardware at continuous full utilization rather than at intermittent or bursty load. Rationale: Ambient and agentic workloads are continuous by nature. Sizing to office-equipment duty cycles produces thermal throttling that is then misdiagnosed as a model limitation. Permalink: https://www.ainativemedical.org/conformance/#anm-4.4 #### ANM-4.5 [MUST] [Class A, B] A conforming deployment MUST provide backup power sufficient to bring inference hardware and storage to an orderly shutdown without loss of tenant data. Permalink: https://www.ainativemedical.org/conformance/#anm-4.5 #### ANM-4.6 [SHOULD] [Class A, B] Inference hardware in a conforming deployment SHOULD be sited to keep end-to-end response latency dominated by computation rather than by network transit. Rationale: The architecture's experiential claim is that machine capability feels adjacent. Latency budget spent on transit is the one cost this siting exists to eliminate. Permalink: https://www.ainativemedical.org/conformance/#anm-4.6 #### ANM-4.7 [MUST] [Class C] A Class C shell MUST document its available power capacity, thermal rejection capacity, floor loading, and cable pathway capacity in terms that permit a prospective tenant to size an enclave against them. Permalink: https://www.ainativemedical.org/conformance/#anm-4.7 #### ANM-4.8 [MUST] [Class A, B] A deployment claiming an interactive latency figure MUST publish a glass-to-photon latency budget that allocates the end-to-end path across sensor capture, local transport, inference execution, and render and scan-out, and MUST state the measured contribution of each stage rather than the inference time alone. Rationale: GPU proximity removes wide-area transit from the path and nothing else. Sensor integration, codec conversion, scheduling jitter, and display scan-out routinely exceed the inference kernel, so a figure derived from locality alone is not a claim about what a clinician perceives. Verification: A published budget table whose stage allocations sum to the claimed total, each traceable to an instrumented measurement method. Permalink: https://www.ainativemedical.org/conformance/#anm-4.8 #### ANM-4.9 [MUST] [Class A, B] Latency claims supporting augmented-reality or other interactive clinical guidance MUST be measured end to end at the display surface under representative clinical load, and MUST be reported at the 99th percentile rather than as a mean or best-case value. Rationale: Procedural error is induced by the slow frames, not the average one. A mean figure captured on an idle node conceals exactly the tail behavior that determines whether an overlay is safe to rely on during instrument manipulation. Verification: Measurement records showing p99 end-to-end latency with concurrent inference workloads active, captured at the panel rather than at kernel exit. Permalink: https://www.ainativemedical.org/conformance/#anm-4.9 #### ANM-4.10 [MUST NOT] [Class A, B, C] An implementation MUST NOT characterize inference latency as sub-millisecond, instantaneous, or real-time when describing an end-to-end interactive path that includes sensor capture and display rendering. Rationale: Sub-millisecond figures describe individual operations such as an in-memory graph traversal. Applying them to a perceptual path that necessarily includes capture and scan-out overstates the system's behavior by an order of magnitude. Permalink: https://www.ainativemedical.org/conformance/#anm-4.10 ### 5. The Acoustic Enclave Physical containment of the captured field. Continuous ambient capture is defensible only where the room can be shown to contain what it hears, which makes acoustics a security control. Formalizes: The Sovereign Enclave: The Architecture of the Hardened Shell (https://www.ainativemedical.org/sections/enclave/) #### ANM-5.1 [MUST] [Class A] An enclave in a Class A deployment MUST achieve a Sound Transmission Class rating of not less than STC 55 across every partition, door, and penetration bounding the captured acoustic field. Rationale: STC 55 is a laboratory assembly rating adopted here as a construction floor, not a statement about speech intelligibility in the delivered room. Stating a number converts confidentiality from an assertion into an inspectable building property; ANM-5.7 states the field criterion the number is intended to achieve. Verification: Field testing of the assembled construction, not the rated assembly specification alone. Permalink: https://www.ainativemedical.org/conformance/#anm-5.1 #### ANM-5.2 [MUST] [Class A] Acoustic performance in a Class A deployment MUST be verified by field measurement of the constructed enclave after installation of all services, and MUST NOT be claimed solely on the basis of laboratory ratings for the specified assemblies. Rationale: Rated assemblies routinely underperform once penetrated by conduit, ductwork, and outlets. The delivered room is the only meaningful subject of the measurement. Permalink: https://www.ainativemedical.org/conformance/#anm-5.2 #### ANM-5.3 [MUST] [Class A] Every mechanical, electrical, and plumbing penetration of a Class A enclave boundary MUST be acoustically sealed and MUST be included in the verification required by ANM-5.2. Permalink: https://www.ainativemedical.org/conformance/#anm-5.3 #### ANM-5.4 [MUST] [Class A] A Class A enclave MUST maintain an ambient noise floor low enough for reliable speech capture at the far field of the room, so that ingestion accuracy does not depend on participants addressing a device directly. Rationale: Ambient intelligence fails quietly when the room is noisy: the system degrades to capturing only the loudest speaker, which is rarely the most consequential one. Permalink: https://www.ainativemedical.org/conformance/#anm-5.4 #### ANM-5.5 [SHOULD] [Class B] A Class B deployment SHOULD apply the acoustic requirements of this chapter to any space in which privileged material is displayed or discussed, notwithstanding the absence of ambient capture. Permalink: https://www.ainativemedical.org/conformance/#anm-5.5 #### ANM-5.6 [MUST] [Class C] A Class C shell claiming acoustic readiness MUST identify which specific spaces are capable of achieving STC 55 and what construction is outstanding, and MUST NOT represent an unbuilt rating as achieved. Permalink: https://www.ainativemedical.org/conformance/#anm-5.6 #### ANM-5.7 [MUST] [Class A] A Class A enclave MUST demonstrate confidential speech privacy in the constructed room by achieving a Privacy Index greater than 95% and an Articulation Index below 0.05, measured under ASTM E1130 field testing conditions. Rationale: Sound Transmission Class rates an assembly in a laboratory; it does not measure the intelligibility of speech leaving a finished room with doors, returns, and flanking paths. Privacy Index and Articulation Index are the recognized field metrics for confidential speech privacy, which makes them the correct basis for a security claim. Verification: An ASTM E1130 field measurement report for each enclave, produced after all services are installed, stating measured PI and AI values against the source and receiver positions used. Permalink: https://www.ainativemedical.org/conformance/#anm-5.7 #### ANM-5.8 [MUST NOT] [Class A, B, C] An implementation MUST NOT represent any Sound Transmission Class rating as rendering speech inaudible, unrecoverable, or immune to reconstruction, and MUST NOT describe an acoustic assembly as an air gap. Rationale: Acoustic isolation reduces intelligibility to a measurable threshold under defined test conditions. It does not defeat an instrumented adversary using structural or vibration-based capture, and stating otherwise misrepresents the control to reviewers who may rely on it. Permalink: https://www.ainativemedical.org/conformance/#anm-5.8 ### 6. Sensory Ingestion How ambient reality enters the system, and what the ingestion layer is forbidden to retain. Statelessness is required here precisely because raw capture is the largest available liability. Formalizes: Stateless Multimodal Routing: The Ingestion Pipeline (https://www.ainativemedical.org/sections/ingestion/) #### ANM-6.1 [MUST NOT] [Class A] The ingestion layer of a Class A deployment MUST NOT persist raw uncompressed acoustic or spatial capture to durable storage at any point in its processing pipeline. Rationale: A durable archive of everything ever said in an institution is an extraordinary liability and an unnecessary one, because the structured product of the capture is what carries the value. Verification: Storage inspection during and after an active capture session shows no raw retention. Permalink: https://www.ainativemedical.org/conformance/#anm-6.1 #### ANM-6.2 [MUST] [Class A] The ingestion layer of a Class A deployment MUST reduce ambient capture to structured records in flight, and MUST discard the source capture once reduction completes. Permalink: https://www.ainativemedical.org/conformance/#anm-6.2 #### ANM-6.3 [MUST] [Class A] A Class A deployment MUST make the operating state of ambient capture perceptible to every person present in the enclave without requiring that person to consult a screen or an application. Rationale: Consent to ambient capture is meaningless if its subjects cannot tell whether it is active. The indication belongs to the room, not to a settings panel. Permalink: https://www.ainativemedical.org/conformance/#anm-6.3 #### ANM-6.4 [MUST] [Class A] A Class A deployment MUST provide an in-room means of suspending ambient capture that is available to any occupant and that takes effect without administrative approval. Permalink: https://www.ainativemedical.org/conformance/#anm-6.4 #### ANM-6.5 [MUST] [Class A] Structured records derived from ambient capture MUST remain inside the demarcation boundary and MUST inherit every prohibition of Chapter 3 that applies to inference payloads. Rationale: Derived records are frequently treated as a different class of data than the capture they came from. They are not, and the boundary must not distinguish them. Permalink: https://www.ainativemedical.org/conformance/#anm-6.5 #### ANM-6.6 [SHOULD] [Class A] A Class A deployment SHOULD support per-session exclusion of identified participants from ambient capture, so that privilege, works-council obligations, and individual objection can be honored without disabling the room. Permalink: https://www.ainativemedical.org/conformance/#anm-6.6 #### ANM-6.7 [MUST NOT] [Class B] A Class B deployment MUST NOT represent itself as providing ambient intelligence, ambient capture, or continuous sensory ingestion. Permalink: https://www.ainativemedical.org/conformance/#anm-6.7 ### 7. Orchestration & Agent Authority The bounds within which autonomous software may act. These clauses constrain tool invocation, including invocation through the Model Context Protocol, to authority that is physically established. Formalizes: Edge-Native Agentic Orchestration: The Orchestration Daemon (https://www.ainativemedical.org/sections/orchestration/) #### ANM-7.1 [MUST] [Class A, B] The orchestration layer of a conforming deployment MUST execute inside the demarcation boundary, including its policy evaluation, routing decisions, and scheduling state. Rationale: An orchestrator hosted outside the boundary observes every request it routes, which reproduces the exposure the boundary was drawn to prevent. Permalink: https://www.ainativemedical.org/conformance/#anm-7.1 #### ANM-7.2 [MUST] [Class A, B] Every tool invocation available to an autonomous agent in a conforming deployment MUST be declared in advance, and an agent MUST NOT acquire a capability at runtime that was not present in its declared set. Rationale: Dynamic capability acquisition makes an agent's authority unbounded and unauditable, which no regulated institution can grant standing access under. Permalink: https://www.ainativemedical.org/conformance/#anm-7.2 #### ANM-7.3 [MUST] [Class A, B] Tool invocation through the Model Context Protocol in a conforming deployment MUST be authorized against the physical identity established under Chapter 8, and MUST be denied when no authorizing presence is established. Rationale: This is the specific point at which the agentic workload meets the physical architecture: an agent's reach is bounded by who is verifiably in the room, not by a credential that may have leaked. Permalink: https://www.ainativemedical.org/conformance/#anm-7.3 #### ANM-7.4 [MUST] [Class A, B] A conforming deployment MUST record every autonomous tool invocation with the invoking agent, the authorizing identity, the parameters supplied, and the outcome, and MUST retain that record inside the demarcation boundary. Permalink: https://www.ainativemedical.org/conformance/#anm-7.4 #### ANM-7.5 [MUST] [Class A, B] A conforming deployment MUST classify tool invocations that mutate external state, transfer value, or communicate outside the organization as requiring explicit human authorization for each occurrence. Rationale: Autonomy is acceptable for reasoning and retrieval and unacceptable for irreversible action. The line is drawn at consequence, not at capability. Permalink: https://www.ainativemedical.org/conformance/#anm-7.5 #### ANM-7.6 [MUST] [Class A, B] Retrieval assets built from tenant material — indexes, knowledge graphs, embeddings, and evaluation sets — MUST be stored inside the demarcation boundary and MUST be owned by the tenant. Permalink: https://www.ainativemedical.org/conformance/#anm-7.6 #### ANM-7.7 [SHOULD] [Class A] A conforming deployment SHOULD express retrieval over typed relationships between people, documents, decisions, and events rather than over undifferentiated similarity alone. Rationale: The ambient record's distinctive value is relational: who met whom, about what, and in what order. Flat similarity search discards precisely that structure. Permalink: https://www.ainativemedical.org/conformance/#anm-7.7 #### ANM-7.8 [MUST] [Class A, B] A conforming deployment MUST interpose a policy and authorization engine between the agent and the Model Context Protocol server, such that no tool invocation reaches building hardware or clinical records on the model's authority alone. Policy MUST be externalized from both the model weights and the Model Context Protocol server, MUST be evaluated deterministically, and MUST default to deny. Rationale: A model that ingests ambient clinical dialogue is ingesting untrusted input. Placing the authorization decision outside the model is what prevents a phrase spoken in the room, dictated from a patient's document, or embedded in a scanned referral from redirecting the agent, and it keeps the decision reviewable by an examiner. Verification: Configuration evidence showing an external decision engine in the invocation path, its policy corpus under version control, and a default-deny result for an undeclared tool. Permalink: https://www.ainativemedical.org/conformance/#anm-7.8 #### ANM-7.9 [MUST] [Class A, B] A conforming deployment MUST treat all ambient capture, patient-supplied documents, and third-party correspondence entering the inference path as untrusted input with respect to agent authority, and MUST NOT allow instructions originating in that content to alter the agent's declared tool set, its authorization scope, or the policy governing it. Rationale: Prompt injection, privilege escalation through chained invocations, and confused-deputy execution against door-strike or record-retrieval interfaces are the governing risks of an agentic clinical deployment. None of them are mitigated by processing the model locally. Verification: Documented adversarial testing against injection and escalation attempts, with results retained inside the boundary. Permalink: https://www.ainativemedical.org/conformance/#anm-7.9 #### ANM-7.10 [MUST] [Class A, B] Transport between the agent, the policy and authorization engine, and the Model Context Protocol server MUST be mutually authenticated using short-lived workload credentials, and authorization MUST be expressed as fine-grained access control enumerated per tool rather than as a single privileged service identity. Rationale: The local network sits inside the demarcation boundary but not inside the trust boundary. Locality is not a substitute for authenticating the workloads that speak to each other across it. Verification: Certificate and policy configuration showing mutual authentication on each hop and a per-tool authorization matrix. Permalink: https://www.ainativemedical.org/conformance/#anm-7.10 #### ANM-7.11 [MUST NOT] [Class A, B, C] An implementation MUST NOT treat a discovery artifact such as an llm.txt file as a source of authority, capability, or policy. Machine-level execution MUST be carried by the Model Context Protocol and local programmatic endpoints, and a conforming deployment MUST operate fully with no discovery artifact present. Rationale: A discovery file is a documentation convention that is trivially editable and unauthenticated. Elevating it to a system-level protocol would place the enclave's capability model in a file that confers no cryptographic assurance whatsoever. Permalink: https://www.ainativemedical.org/conformance/#anm-7.11 ### 8. Identity & Physical Access Entry to the enclave is an authentication event. These clauses require that physical presence be established, recorded, and bound to the inference sessions it authorizes. Formalizes: Zero-Trust Physical Identity & The MCP Standard (https://www.ainativemedical.org/sections/identity/) #### ANM-8.1 [MUST] [Class A, B] A conforming deployment MUST treat entry to the enclave as an authentication event of equal standing to a software credential, and MUST record it as such. Rationale: A sovereign compute environment is only as strong as its physical access log. Treating the door as facilities management rather than as identity infrastructure leaves the strongest control unrecorded. Permalink: https://www.ainativemedical.org/conformance/#anm-8.1 #### ANM-8.2 [MUST] [Class A, B] A conforming deployment MUST bind each inference session to the physical identity or identities established as present in the enclave at the time the session is initiated. Permalink: https://www.ainativemedical.org/conformance/#anm-8.2 #### ANM-8.3 [MUST] [Class A, B] Physical access records for a conforming deployment MUST be retained inside the demarcation boundary and MUST be subject to the prohibitions of Chapter 3. Rationale: Access logs describe who was in the room and when, which is itself privileged information in a transaction, litigation, or clinical context. Permalink: https://www.ainativemedical.org/conformance/#anm-8.3 #### ANM-8.4 [MUST NOT] [Class A, B] A conforming deployment MUST NOT permit administrative access to inference hardware, storage, or orchestration state from outside its demarcation boundary. Rationale: Remote administrative access is a payload-capable path with the highest privilege in the system, and its convenience is the most common reason sovereignty claims fail on inspection. Permalink: https://www.ainativemedical.org/conformance/#anm-8.4 #### ANM-8.5 [MUST] [Class A, B] Maintenance performed by a software integrator MUST occur under an identity distinct from any tenant identity, and MUST be recorded with the same fidelity required of tenant access by ANM-8.1. Permalink: https://www.ainativemedical.org/conformance/#anm-8.5 #### ANM-8.6 [SHOULD] [Class A] A conforming deployment SHOULD detect and record the presence of unenrolled individuals in the enclave during an active ambient capture session. Permalink: https://www.ainativemedical.org/conformance/#anm-8.6 #### ANM-8.7 [MUST] [Class C] A Class C shell MUST document the physical access control provisions available at the intended enclave location, and MUST NOT claim identity binding, because no inference sessions exist to bind. Permalink: https://www.ainativemedical.org/conformance/#anm-8.7 ### 9. Ownership & Governance The Tripartite Ownership Model, stated as enforceable separations rather than as commercial preference. These clauses define what each party is forbidden to hold. Formalizes: Real Estate Economics & Landlord-Tenant Alignment (https://www.ainativemedical.org/sections/economics/) #### ANM-9.1 [MUST] [Class A, B] A conforming deployment MUST separate the property owner, the tenant, and the software integrator into distinct parties whose holdings do not overlap, in accordance with the Tripartite Ownership Model. Verification: Executed agreements reflect the separation and are available for examination. Permalink: https://www.ainativemedical.org/conformance/#anm-9.1 #### ANM-9.2 [MUST NOT] [Class A, B] The property owner in a conforming deployment MUST NOT hold ownership of, access to, or a contingent interest in tenant compute hardware, inference data, retrieval assets, or model outputs. Rationale: This separation is what allows a landlord to finance and install sovereign infrastructure without acquiring rights that would make the tenant's custody claim unsustainable. Permalink: https://www.ainativemedical.org/conformance/#anm-9.2 #### ANM-9.3 [MUST NOT] [Class A, B] The software integrator in a conforming deployment MUST NOT hold ownership of tenant data, policies, evaluations, routing logic, retrieval assets, or commissioned model adaptations. Permalink: https://www.ainativemedical.org/conformance/#anm-9.3 #### ANM-9.4 [MUST] [Class A, B] A conforming deployment MUST provide the tenant with a documented exit under which inference capability, retrieval assets, and accumulated institutional memory remain operable after termination of any agreement with the software integrator or the property owner. Rationale: Sovereignty that evaporates at contract termination was vendor dependence with a longer notice period. Permalink: https://www.ainativemedical.org/conformance/#anm-9.4 #### ANM-9.5 [MUST] [Class A, B] A conforming deployment MUST disclose to the tenant every third-party license, model license, and usage restriction that constrains the tenant's use of outputs produced within the enclave. Permalink: https://www.ainativemedical.org/conformance/#anm-9.5 #### ANM-9.6 [MUST NOT] [Class A, B] A conforming deployment MUST NOT use tenant material to train, fine-tune, evaluate, or improve any model or system made available to another party. Rationale: Cross-tenant improvement is the mechanism by which a sovereignty claim is most often quietly voided, and it is rarely visible in the operating architecture. Permalink: https://www.ainativemedical.org/conformance/#anm-9.6 #### ANM-9.7 [MUST] [Class C] A Class C shell MUST disclose the ownership structure under which a future enclave would be delivered, so that a prospective tenant can evaluate the separation required by ANM-9.1 before committing. Permalink: https://www.ainativemedical.org/conformance/#anm-9.7 #### ANM-9.8 [MUST] [Class A, B, C] Where the property owner is or may be a referral source for the tenant, space and compute arrangements between the parties MUST satisfy an applicable rental exception under 42 CFR § 411.357 in full, including a signed written agreement, a term of at least one year, a description of the premises and equipment covered, space and equipment not exceeding what is reasonable and necessary for the tenant's legitimate business purposes, compensation set in advance at fair market value, and commercial reasonableness assessed independent of referrals. Rationale: Fair market value is a necessary condition of a defensible arrangement, not a safe harbor in itself. The rental exceptions impose several further conditions, and an arrangement priced correctly but structured loosely still fails. Verification: Executed agreements together with a contemporaneous independent valuation, refreshed on a defined cycle rather than performed once at signing. Permalink: https://www.ainativemedical.org/conformance/#anm-9.8 #### ANM-9.9 [MUST NOT] [Class A, B, C] Compute pricing, capacity allocation, tiering, escalation, discounts, and service credits MUST NOT be determined in any manner that takes into account the volume or value of referrals or other business generated between the parties. Percentage-of-revenue and per-referral compute pricing MUST NOT be used. Rationale: Capacity provisioned by reference to a tenant's referral footprint rather than its clinical throughput is the specific failure mode this architecture could otherwise enable at scale, and the Anti-Kickback Statute turns on intent rather than on valuation mechanics. Verification: The compute rate schedule and allocation methodology, documented in terms of throughput and capacity units with no referral-derived variable. Permalink: https://www.ainativemedical.org/conformance/#anm-9.9 ### 10. Auditability & Evidence What a deployment must be able to show an examiner. The specification's compliance argument is structural, which obligates it to be demonstrable on inspection. Formalizes: The Compliance Moat (https://www.ainativemedical.org/sections/compliance/) #### ANM-10.1 [MUST] [Class A, B] A conforming deployment MUST be able to demonstrate the absence of an egress path for inference payloads to an examiner on site, without relying on an attestation issued by a third party. Rationale: The specification's compliance argument is that architecture can be shown rather than asserted. That claim obligates the deployment to be demonstrable on inspection. Permalink: https://www.ainativemedical.org/conformance/#anm-10.1 #### ANM-10.2 [MUST] [Class A, B] A conforming deployment MUST maintain records of physical access, tool invocation, model and software version history, and boundary configuration changes, sufficient to reconstruct the operating state of the enclave at any past point within its retention period. Permalink: https://www.ainativemedical.org/conformance/#anm-10.2 #### ANM-10.3 [MUST] [Class A, B] Audit records in a conforming deployment MUST be append-only and MUST NOT be alterable by the software integrator. Rationale: An audit record that the operating party can edit does not constrain the operating party. Permalink: https://www.ainativemedical.org/conformance/#anm-10.3 #### ANM-10.4 [MUST] [Class A, B] A conforming deployment MUST identify the specific statutory or regulatory obligations its architecture is intended to satisfy, and MUST map each to the clauses of this specification relied upon. Rationale: A structural compliance claim is only useful if it names what it is compliant with. An unmapped claim cannot be examined and should not be credited. Permalink: https://www.ainativemedical.org/conformance/#anm-10.4 #### ANM-10.5 [SHOULD] [Class A, B] A conforming deployment SHOULD re-verify the acoustic performance required by Chapter 5 and the path enumeration required by ANM-2.2 after any construction, reconfiguration, or hardware change affecting the enclave. Permalink: https://www.ainativemedical.org/conformance/#anm-10.5 #### ANM-10.6 [MUST] [Class A, B, C] A conforming deployment MUST publish the date of its most recent conformance self-assessment alongside any conformance claim it makes. Permalink: https://www.ainativemedical.org/conformance/#anm-10.6 --- # Glossary > Canonical definitions of the terms used in the AI-Native Medical specification: sovereign compute edge node, zero egress, the Tripartite Ownership Model, ambient intelligence, and the acoustic and identity requirements they depend on. ## Core concepts ### AI-Native Medical The AI-Native Medical Office Building is a healthcare real estate asset engineered so that the building itself performs clinical inference: a sovereign, on-premises compute edge node in which GPU hardware, acoustic isolation, ambient sensing, and physical identity enforcement are delivered as building infrastructure rather than as a cloud subscription. Protected Health Information is ingested and processed locally within a bounded glass-to-photon latency budget, and no inference payload crosses the property's network boundary. Zero egress operates as a physical and technical control within a defense-in-depth HIPAA Security Architecture pursuant to 45 CFR § 164.312 — narrowing the Business Associate Agreement surface to direct local infrastructure operators and making the control inspectable in the real estate itself, rather than establishing regulatory compliance by locality alone. Also written: AI native medical; AI-Native Healthcare; The Exam Room as the Machine This is a vendor-neutral specification describing a class of physical infrastructure, not a software product and not a particular building. Its unit of delivery is a leasable enclave, which makes commercial real estate — rather than the hyperscaler — the vehicle through which regulated enterprises obtain frontier AI capability. Permalink: https://www.ainativemedical.org/glossary/#ai-native-office ### Ambient Clinical AI Ambient clinical AI is a software condition in which AI agents operate as persistent participants in care — listening to encounters, drafting documentation, reconciling data across the record, and invoking downstream systems autonomously. The AI-Native Medical specification treats ambient clinical AI as a workload description rather than an architecture: it names what the software does, not where the computation physically occurs or who holds custody of the protected health information. Also written: agentic clinic; ambient clinical intelligence; AI clinical agents; ambient documentation The distinction is load-bearing. Persistent, autonomous clinical agents generate continuous inference against a covered entity's most sensitive material, which is precisely the access pattern that metered-egress cloud infrastructure prices punitively and that HIPAA-covered institutions cannot lawfully authorize off-premises. Ambient clinical AI is therefore the demand; the AI-Native Medical is the substrate that demand requires. Permalink: https://www.ainativemedical.org/glossary/#agentic-office ### Ambient Intelligence Ambient intelligence, in the AI-Native Medical, is machine capability that engages the work without being invoked. Rather than waiting on a prompt typed into a chat interface, the environment continuously perceives the meeting, the document, and the decision as they occur. The specification treats the keyboard as a legacy ingestion bottleneck and the room as a continuous sensory organ that supersedes it. Also written: ambient AI; ambient computing; ambient telemetry Permalink: https://www.ainativemedical.org/glossary/#ambient-intelligence ### Intelligence Compounding Intelligence compounding is the accrual effect the AI-Native Medical is designed to capture: because ambient context is structured and retained inside the tenant's own boundary, each engagement improves the retrieval assets available to the next. The resulting institutional memory is an owned asset rather than a vendor-held one, and it cannot be replicated by a competitor or withdrawn at renewal. Also written: compounding intelligence; institutional memory; the flywheel Permalink: https://www.ainativemedical.org/glossary/#intelligence-compounding ### Structural Ceiling The structural ceiling is the problem the AI-Native Medical exists to solve: the institutions with the most to gain from frontier AI — banks, law firms, healthcare systems — are the ones least able to adopt it as delivered, because their regulatory obligations forbid the data movement that cloud inference requires. The limit is architectural, not a matter of budget, appetite, or technical sophistication. Also written: the structural ceiling; regulated AI ceiling Permalink: https://www.ainativemedical.org/glossary/#structural-ceiling ## Architecture ### Sovereign Compute Edge Node A sovereign compute edge node is the deployable unit of the AI-Native Medical: localized accelerator hardware, storage, and orchestration sited inside a tenant-controlled physical boundary, where the tenant owns the silicon, the inference data, and the model outputs outright. Because inference executes adjacent to the people generating the work, latency is a function of distance measured in meters rather than network topology. Also written: sovereign compute; sovereign edge node; on-premises inference node Permalink: https://www.ainativemedical.org/glossary/#sovereign-compute-edge-node ### Demarcation Boundary The demarcation boundary is the explicitly declared physical and logical perimeter of an AI-Native Medical deployment, inside which the tenant holds sole custody of data and compute. It is the surface against which the specification's zero-egress requirement is evaluated: a conforming deployment must be able to enumerate every path that crosses the boundary and demonstrate that no inference payload traverses any of them. Also written: tenant boundary; sovereignty boundary; declared boundary Permalink: https://www.ainativemedical.org/glossary/#demarcation-boundary ### Stateless Ingestion Stateless ingestion is the AI-Native Medical's requirement that the layer converting ambient reality into structured data retain no raw capture. Acoustic and spatial input is transcribed and reduced to lightweight structured records in flight; the raw uncompressed capture is never written to durable storage. The AI-Native Medical adopts this constraint because a durable store of raw ambient recording is both a liability surface and an unnecessary one. Also written: stateless ingestion layer; ephemeral ingestion Statelessness is a property of the ingestion layer, not of the system as a whole. The structured artifacts ingestion produces — clinical notes, triage scores, finalized reports, and the audit records evidencing them — are deliberately retained, and become persistent Protected Health Information subject to the full Security Rule safeguard set, the practice's retention schedule, and breach-notification obligations once written or filed to the record. Permalink: https://www.ainativemedical.org/glossary/#stateless-ingestion ### Localized Orchestration Layer The localized orchestration layer is the AI-Native Medical's central logic unit: it functions as a hypervisor for physical space, abstracting the room's sensory hardware into resources that software workloads can schedule against. Once ambient input has been routed, transcribed, and structured, the orchestration layer is what evaluates it against policy and triggers autonomous action inside the demarcation boundary. Also written: orchestration layer; hypervisor for physical space; spatial hypervisor Permalink: https://www.ainativemedical.org/glossary/#orchestration-layer ### Model Context Protocol (MCP) The Model Context Protocol is the open standard by which models invoke tools and reach external context, and it is the interface through which agents act inside an AI-Native Medical. The specification's contribution is physical: it requires MCP tool invocation to be gated by zero-trust physical identity, so an agent's authority is bounded by who is verifiably present in the enclave. Also written: MCP; MCP server; Model Context Protocol server This is the point at which the AI-Native Medical and the agentic office meet concretely. MCP describes what an agent may call; the AI-Native Medical constrains where that call may execute and whose presence authorizes it. Because a model ingesting ambient clinical dialogue is ingesting untrusted input, the specification requires a Policy & Authorization Engine — Open Policy Agent or equivalent, over mTLS with fine-grained role-based access control — to sit between the agent and the MCP server and adjudicate every tool call under default-deny policy, which is what bounds prompt injection and privilege escalation to a decision the enclave makes rather than one the model makes. Permalink: https://www.ainativemedical.org/glossary/#model-context-protocol ### GraphRAG GraphRAG is retrieval-augmented generation over an explicit knowledge graph rather than over flat vector similarity alone, and it is the retrieval strategy the AI-Native Medical uses to exploit ambient context. Because the AI-Native Medical observes who met with whom, about what, and in what sequence, it can build typed relationships between people, documents, and decisions that similarity search alone cannot recover. Also written: graph retrieval augmented generation; graph RAG; knowledge graph retrieval Permalink: https://www.ainativemedical.org/glossary/#graphrag ## Data & economics ### Zero Egress Zero egress is the defining data-movement property of the AI-Native Medical: no inference payload, no ambient telemetry, and no derived artifact crosses the tenant's demarcation boundary during normal operation. Because sensitive material never transits a third-party network, the AI-Native Medical removes the vendor-trust dependency that procedural cloud compliance is built to manage, and eliminates per-gigabyte egress billing entirely. Also written: zero-egress architecture; no data egress; egress-free inference Zero egress is stated as an architectural constraint, not a configuration setting or a contractual promise. In the AI-Native Medical there is no supported path by which inference data leaves the premises. It functions as a physical and technical control within a defense-in-depth HIPAA Security Architecture under 45 CFR § 164.312 — narrowing the Business Associate Agreement surface to direct local operators — and does not by itself discharge the administrative, physical, and technical safeguards the Security Rule requires. Permalink: https://www.ainativemedical.org/glossary/#zero-egress ### Egress Asymmetry Egress asymmetry is the hyperscaler pricing structure the AI-Native Medical is designed to escape: inbound data transfer is free or subsidized while outbound transfer is metered and billed. The asymmetry is not incidental. It makes accumulated corporate data progressively more expensive to relocate the more of it exists, converting a storage relationship into a structural dependency the specification refers to as data gravity. Also written: asymmetric egress pricing; data gravity; egress economics Permalink: https://www.ainativemedical.org/glossary/#egress-asymmetry ## Physical environment ### Acoustic Enclave The acoustic enclave is the engineered room in which an AI-Native Medical deployment operates. Because continuous ambient capture is only defensible if the captured field is contained, the specification treats acoustic isolation as a security control rather than a comfort amenity, and specifies it against measurable confidential-speech-privacy criteria: STC-55 assemblies engineered to a Privacy Index above 95% and an Articulation Index below 0.05, verified by ASTM E1130 field testing of the constructed room. Also written: the enclave; acoustic isolation; sovereign enclave Permalink: https://www.ainativemedical.org/glossary/#acoustic-enclave ### STC 55 STC 55 is the Sound Transmission Class rating the AI-Native Medical specifies for enclave partitions. Sound Transmission Class measures how effectively an assembly attenuates airborne sound in a laboratory setting. The AI-Native Medical adopts STC 55 as a numeric floor, then requires the delivered room to demonstrate confidential speech privacy in the field — a Privacy Index above 95% and an Articulation Index below 0.05 under ASTM E1130 — because a laboratory rating is not by itself a statement about speech intelligibility at the boundary. Also written: Sound Transmission Class 55; STC rating; STC-55 partition Permalink: https://www.ainativemedical.org/glossary/#stc-55 ## Governance & compliance ### Tripartite Ownership Model The Tripartite Ownership Model is the AI-Native Medical's governance architecture, separating a deployment into three parties with strictly disjoint holdings: the property owner, who provides the shell and physical infrastructure; the tenant, who owns the compute hardware, the inference data, and all outputs; and the software operator, who maintains orchestration inside the boundary. No party holds access to what belongs to another. Also written: tripartite ownership; three-party ownership model The Tripartite Ownership Model is what makes the AI-Native Medical commercially deliverable through a lease. It lets a landlord finance and install sovereign infrastructure without ever acquiring rights to tenant data, and it lets a software operator maintain the stack without custody of what it operates on. Permalink: https://www.ainativemedical.org/glossary/#tripartite-ownership-model ### Physical Sovereignty Physical sovereignty is the AI-Native Medical's claim that control over data follows from physical custody of the hardware processing it. Where cloud sovereignty is asserted through contract, jurisdiction, and vendor attestation, the AI-Native Medical grounds it in a locked room containing tenant-owned silicon — a control an auditor can walk into, inspect, and verify without relying on a third party's representation. Also written: data sovereignty; sovereignty by architecture Permalink: https://www.ainativemedical.org/glossary/#physical-sovereignty ### Zero-Trust Physical Identity Zero-trust physical identity is the AI-Native Medical's requirement that entry to the enclave be treated as an authentication event of equal standing to a software credential. A sovereign compute environment is only as strong as its physical access log, so the specification binds identity at the door to identity in the inference session — presence in the room is itself an authorization fact, recorded and auditable. Also written: physical identity; physical access identity; zero trust physical layer Permalink: https://www.ainativemedical.org/glossary/#zero-trust-physical-identity ### Compliance Moat The compliance moat is the competitive position the AI-Native Medical creates by making its central control structural rather than procedural. Cloud compliance rests on access controls and audit logs maintained by a third party whose interests are not identical to the customer's; the AI-Native Medical instead makes third-party data movement physically unavailable, so that control is demonstrated by inspection rather than attested by process. The moat is the durability of that one control, not a claim that the architecture completes a compliance program. Also written: architecture as compliance; structural compliance Permalink: https://www.ainativemedical.org/glossary/#compliance-moat ### Software Integrator The software integrator is the third party in the AI-Native Medical's Tripartite Ownership Model: the operator that deploys, configures, observes, and updates the orchestration stack inside the tenant's demarcation boundary. The role is defined by what it does not hold — the software integrator has no ownership of tenant data, policies, evaluations, routing logic, or retrieval assets, only responsibility for the machinery acting on them. Also written: software operator; integration layer Permalink: https://www.ainativemedical.org/glossary/#software-integrator Terms defined: 20 --- # Cite This Specification > Canonical citation formats for the AI-Native Medical specification: BibTeX, APA, Chicago, IEEE, and CITATION.cff, with permanent clause-level and version-level identifiers. ## BibTeX @techreport{ainativemedical2026, title = {{AI-Native Medical: Sovereign Clinical AI Edge Infrastructure}}, author = {Timothy Walsh and Parham Alizadeh}, institution = {AI-Native Medical}, type = {Draft Specification (RFC)}, number = {v0.2.0}, year = {2026}, url = {https://www.ainativemedical.org/}, doi = {10.5281/zenodo.21852364}, note = {Revised 2026-08-08} } ## APA 7th edition Walsh, T., & Alizadeh, P. (2026). AI-Native Medical: Sovereign Clinical AI Edge Infrastructure (Version 0.2.0) [Draft specification]. AI-Native Medical. https://doi.org/10.5281/zenodo.21852364 ## Chicago 17th edition Walsh, Timothy, and Parham Alizadeh. "AI-Native Medical: Sovereign Clinical AI Edge Infrastructure." Version 0.2.0. Draft specification. AI-Native Medical, 2026. https://doi.org/10.5281/zenodo.21852364 ## IEEE T. Walsh and P. Alizadeh, "AI-Native Medical: Sovereign Clinical AI Edge Infrastructure," AI-Native Medical, Draft Specification v0.2.0, 2026. [Online]. Available: https://www.ainativemedical.org/ ## CITATION.cff cff-version: 1.2.0 message: "If you reference this specification, please cite it as below." title: "AI-Native Medical: Sovereign Clinical AI Edge Infrastructure" version: "0.2.0" date-released: "2026-08-08" url: "https://www.ainativemedical.org/" repository-code: "https://github.com/ainativeoffice/ai-native-medical-spec" doi: "10.5281/zenodo.21852364" type: software license: CC-BY-4.0 keywords: - AI-native medical - sovereign compute - on-premises inference - zero egress - data sovereignty - edge inference - ambient clinical intelligence - healthcare real estate - medical office building - HIPAA - protected health information - Model Context Protocol abstract: >- The delivery of modern healthcare across suburban and retail Medical Office Buildings faces an operational bottleneck: the friction between cloud-dependent AI platforms and strict regulatory, network, and performance constraints. Contemporary clinical workflows demand continuous multimodal AI — real-time ambient transcription, immediate documentation synthesis, automated triage of high-resolution imaging, and low-latency augmented-reality guidance inside a bounded end-to-end budget — yet routing Protected Health Information across public networks to hyperscale clouds introduces prohibitive latency, recurring egress penalties, an expanded Business Associate Agreement and subprocessor surface, and network saturation risk. This specification defines a new healthcare real estate asset class: a zero-egress, on-premise sovereign edge compute node built into the physical Medical Office Building, powered by tenant-owned NVIDIA L40S-class GPU clusters in acoustically isolated, liquid-cooled utility enclaves. Ambient clinical data is ingested and processed locally — never crossing the property's network boundary — delivering tenant custody of data, elimination of recurring egress cost, and a measured glass-to-photon latency budget within 13 milliseconds for interactive guidance. The specification positions zero egress as a physical and technical control within a defense-in-depth HIPAA Security Architecture pursuant to 45 CFR § 164.312, not as a substitute for the administrative, physical, and technical safeguards that rule requires. Architecture reduces the exposure surface; it does not by itself establish regulatory compliance. authors: - given-names: "Timothy" family-names: "Walsh" email: "tw@ainativemedical.org" affiliation: "TruCast" - given-names: "Parham" family-names: "Alizadeh" email: "parham@ainativemedical.org" affiliation: "North Castle Ventures" identifiers: - type: doi value: "10.5281/zenodo.21852364" description: "Concept DOI — always resolves to the latest version."