Waggle: A Latent Language Program for Machine Agents
A decoder-first program for testing whether machines can coordinate with less prose without surrendering fidelity, authority, or human control.
Abstract
Language-model systems commonly coordinate through natural-language tokens even when both endpoints are machines. Prior work reports that, on particular models and benchmarks, direct transfer of embeddings, activations, hidden states, or KV caches can improve task results or reduce generated tokens and latency. Those findings establish a research direction, not a universal law: compatibility is narrow, reported gains are benchmark-dependent, and recent work shows that supposed latent “superposition” can collapse under common fine-tuning regimes.
We propose Waggle, a staged coordination program, and Kea, the independent decipher and audit system that constrains it. Kea comes first. Waggle begins with a deterministic, content-addressed semantic protocol; only after paired economics and fidelity gates may it advance to learned vectors, a phase-encoded consensus experiment, or classical shared state. Consequential authority never resides in an opaque payload or in Kea’s prose interpretation.
The current implementation remains smaller than the proposal, but it is no longer limited to the original deterministic fixture. A real Apple M4 Max Metal kernel has composed source-separated complex states under independent CPU-side Kea recomputation. In C15d-G1, one frozen Qwen3-14B native prefix was restored once and reused across six branches; every informed control completed all six tasks exactly. Resident-native reuse beat repeated resident full-text reconstruction from branch 3 onward, but it did not beat cached-prefix text or cache-warmed fresh-native processes. The mechanism is real; overall hardware efficiency, token or credit savings, learned language, production safety, and commercial value remain unproven.
The Coordination Question
A transformer’s intermediate computation is continuous, while its public language interface is usually discrete. That fact alone does not prove that a vector “contains more meaning” than a token, or that bypassing tokens will improve a system. A token identifier’s nominal bit width is not semantic channel capacity; neither is a tensor’s element count. Meaning depends on a trained encoder, a receiver, context, noise, and the task used to evaluate the transfer.
What can be measured is narrower and more useful. Natural-language coordination requires autoregressive generation, often repeats shared context, and may add acknowledgements or synthesis that do not improve the accepted artifact. Direct-communication research has reported gains on selected tasks: CIPHER sends an expectation over output embeddings[4]; Cache-to-Cache projects and fuses KV caches[5]; activation communication combines intermediate states[6]; LatentMAS and Interlat explore continuous collaboration and compressed latent transfer[7][8]. These are model- and benchmark-specific results, mostly in research settings.
The motivating KE Studios failure was operational: an earlier multi-agent system created more coordination traffic and synthesis than its useful work justified. That incident establishes a cost problem inside one system, not a general theorem about language. The first remedy is therefore scheduling and protocol discipline—solo-first execution, dependency-ready leaves, shared references, bounded retries—not a learned secret language.
The goal is not to make machines harder to understand. It is to stop paying for unnecessary prose while making every consequential state change easier to inspect.
The research question is: under matched models, tasks, tools, authority, and stopping rules, can a negotiated machine transport reduce total coordination cost or latency without reducing accepted quality, policy parity, or human control after Kea’s full overhead is counted? A “yes” must be demonstrated per compatibility domain. A failure at one rung does not license a larger model or a looser metric; it kills or narrows that rung until a new falsifiable mechanism is registered.
Claims and Boundaries
Waggle separates four things that agent stacks often blur: stable meaning, payload encoding, human interpretation, and execution authority.
| Plane | Purpose | Invariant |
|---|---|---|
| Semantic control | Mission, sender, receiver, class, causal parents, references, codec, sensitivity, requested authority, and idempotency | Deterministic and inspectable |
| Machine transport | Structured delta, symbolic packet, learned vector, compatible cache, or declared fallback | No silent codec guessing |
| Kea interpretation | Gloss, claims, alternatives, uncertainty, verification, anomaly signals, and replay | Never replaces the raw record |
| Capability execution | Governed operation with authority lease, budget, side effects, artifact, and evidence | Opaque payloads and glosses cannot authorize work |
Closed model APIs generally do not expose hidden states or KV caches. Waggle therefore does not assume that GPT, Claude, Grok, or any other closed runtime shares a latent dialect. Such systems use the deterministic semantic packet or structured-text fallback unless an official interface explicitly supports another transport. Open runtimes require a declared model, layer, tokenizer, adapter, codec, and decoder compatibility domain.
The word superposition is used operationally, not physically: a learned representation may preserve evidence about more than one registered candidate hypothesis. Toy-model interpretability work demonstrates feature superposition in constrained neural networks, not a general capacity theorem for production LLM communication[29]. No semantic capacity is inferred from vector dimension alone.
Evidence and Prior Art
The v0.3 review searched five lanes: latent inter-model transport, continuous reasoning, emergent-language translation, vector-symbolic and phase models, and commercial adjacency. A 2026 survey independently organizes 18 representative latent-communication methods and identifies compatibility, security, compression, and latent-reasoning questions as open challenges[9]. The dated queries, inclusion rules, exclusions, and claim-strength rubric are recorded in RESEARCH-METHOD.md. This is a bounded landscape review, not a patent search or proof of priority.
What the literature supports
Continuous reasoning can feed hidden states back into a model and can exhibit parallel-search behavior in specific constructions and tasks[1][2]. Direct inter-model methods can outperform prose baselines in their reported setups: CIPHER reports 0.5–5.0 percentage-point gains across five reasoning tasks; the current Cache-to-Cache revision reports an average 2.5× latency speedup and approximately 3.1–5.4% improvement over its text-communication comparison; activation communication reports up to 27% improvement with less than one-quarter of the compute in its experiments; LatentMAS reports 70.8–83.7% fewer output tokens and up to 14.6% higher accuracy across its benchmark suite[4][5][6][7]. These are the source papers’ results, not Waggle results.
Deciphering machine communication also has prior art. Translating Neuralese evaluates translations by their effect on a listener’s beliefs and reward[22]; later work translates emergent communication without parallel language data in referential-game settings[23]; Hyperdimensional Probe extracts concepts from LLM representations[24]. Thought Communication proves shared/private latent identifiability in a stated nonparametric latent-variable setting and demonstrates a framework under those assumptions[10]. None of these results proves that an arbitrary hidden state can be faithfully rendered into English.
What threatens the thesis
The strongest counterresult is not a footnote. The Illusion of Superposition finds signs of superposition only in models trained from scratch; in its training-free and fine-tuned regimes, representations collapse or go unused[3]. Because Waggle v1 initially proposes adapters or fine-tuning, this result can kill that design. The response is an experiment, not a rhetorical defense.
Latent channels also move attacks outside visible-text inspection. A 2026 preprint shows attack-associated effects reactivated through latent interventions, particularly KV-cache handoffs[11]. LCGuard reduces reconstruction-based leakage with adversarial transformations but does not solve general intent, drift, or covert-channel detection[12]. Secret-collusion work demonstrates the relevance of steganographic evaluations while reporting limited capabilities in then-current models[28].
Adjacent, not equivalent
Complex vector-symbolic architectures provide binding, bundling, and compositional operations[14][15][16]. AKOrN studies synchronization dynamics inside artificial neurons[13], and PRISM reports controlled phase structure in a 33M-parameter complex-valued language encoder while explicitly declining to generalize to larger scales[17]. These works motivate a phase experiment; they do not demonstrate inter-agent consensus.
Commercial investment validates adjacent demand, not Waggle. Tensormesh reports $24.5 million in total funding for caching-accelerated inference infrastructure[25]; The Token Company is a YC W26 prompt-compression company[26]; Multiverse reports a $215 million round for quantum-inspired model compression[27]. None is evidence that a latent agent language or Kea is commercially validated.
Optical context compression is excluded from the core ladder. DeepSeek-OCR demonstrates text reconstruction from fewer vision tokens[32], while a direct evaluation finds simple learned or pooled encoders match or outperform the image detour[33]. That is a warning against confusing an evocative representation with a useful one.
Design Principles
The registry, raw record, replay, exact v0 decoder, uncertainty contract, capacity field, and independent viewer exist before a Waggle packet may be treated as accepted traffic. A learned codec cannot be trained first and “made interpretable” later.
Mission identity, causal order, references, sensitivity, authority request, and idempotency remain deterministic. Only the payload encoding changes. Incompatible systems fall back explicitly.
Kea may verify an exact v0 object or qualify a learned interpretation. It may abstain. No gloss—however confident—can create an authority lease or execute a capability.
Generated tokens are one line item. Evaluation includes input and cached input, local/GPU compute, memory, latency, Kea decode and verification, retries, synthesis, human intervention, and accepted artifact quality.
A rung that misses its registered fidelity, policy, quality, safety, or total-economics gate stops. More compute is not the automatic answer; a new run requires a new mechanism and protocol version.
A correction creates a new attributable record and causal message. It does not rewrite the raw message or the interpretation that was shown at the time.
Architecture: The Waggle Ladder
Kea is the foundation, not phase two. The current fixture shell exists before v0 benchmark traffic; every later rung inherits its raw-record, replay, fallback, and authority invariants.
v0 — Designed semantic protocol
Typed compact packets refer to content-addressed Context Packs, artifacts, and evidence instead of restating them. v0 can be decoded exactly into the canonical object because its encoder and decoder are deterministic. This proves object reconstruction, not that a short English gloss captures human intent, and not that v0 is cheaper than a fair prose protocol.
v1 — Learned vector transport
Within a declared open-runtime compatibility domain, a sender encoder emits a bounded sequence of continuous vectors; a receiver adapter consumes them; a separately evaluated Kea decoder produces claims, alternatives, confidence, and abstention. Coconut motivates continuous reasoning inside one model[1], but does not establish this inter-agent channel. v1 remains unbuilt.
v1.5 — Phase experiment
Registered hypotheses receive complex coefficients. Composition measures whether independently produced coefficients align or cancel under a fixed operator. AKOrN, FHRR, and PRISM motivate the mechanics, not the outcome[13][14][17]. “Consensus by interference” remains a falsifiable hypothesis.
v2 — Classical shared state
Compatible agents may update a versioned shared working state through attributable deltas. Every snapshot and mutation remains replayable. This may reduce retransmission, but it is ordinary shared memory—not entanglement and not correlation without a physical communication path.
Kea: Decipher Before Traffic
Kea is specified to be independently usable through a registry, raw record, replay API, CLI, viewer, and event surface. The fixture currently provides the registry, raw record, replay, API, CLI, and viewer—not a general event stream. Director may become its richest visual client, but a Director outage must not remove the audit trail.
| Duty | v0 meaning | Learned-rung requirement |
|---|---|---|
| Render | Reconstruct the canonical packet exactly; produce a deterministic gloss | Report calibrated claims, alternatives, OOD state, and abstention—not “the model’s true thought” |
| Log and replay | Preserve raw hash, causal order, codec/decoder versions, interpretation, and correction lineage | Retain the interpretation shown at the time and permit later decoder comparison without rewriting history |
| Watch | Reject unknown codec, invalid schema, oversize payload, causal gap, and failed exact reconstruction | Evaluate distribution drift, adversarial directions, disagreement, and task-context mismatch |
| Decompose | Identify declared required fields versus excess fields | Attempt shared/private or required/excess inference only where validated; otherwise mark unknown |
| Budget | Bound payload and undecodable bytes | Bound operational hidden capacity with a registered red-team benchmark |
| Correct | Append an attributable correction | Preserve alternatives, disagreement, and immutable decoder history |
The fixture alpha currently implements only the left-hand column’s minimum local slice. A self-hash is not a signed registry. A hash chain is not an externally anchored immutable ledger. A hardcoded policy result is not a policy-parity test. These are named gaps, not hidden implementation details.
Emergent-language translation and probing research show that useful machine-message interpretation can be evaluated[22][23][24]. They do not guarantee a universal decoder. Kea is therefore a measurement program, not a promise of perfect access to latent meaning or private chain-of-thought.
Formal Mechanisms
This section defines what the headline mechanisms mean well enough to fail. Symbols describe the proposed experiment; they are not evidence that the mechanism works.
7.1 Canonical semantic packet
Let a semantic packet be s = (c, p, r, h), where c is deterministic control metadata, p is a typed payload, r is the set of immutable context/artifact/evidence references, and h is the payload and envelope integrity record. Authority is deliberately absent; the packet may contain only none, proposal, or request.
For v0, the pass criterion is byte-canonical object equality after encode/decode. This is a protocol property, not semantic equivalence between arbitrary prose and a human’s intended meaning.
7.2 Learned channel
For a registered compatibility domain, the sender produces z = Eθ(s, x) in a bounded k × d vector channel. The receiver returns task output ŷ = Rφ(z, xᵣ). Kea independently returns an interpretation distribution q = Kψ(z, c) plus an abstention decision.
Training loss is not the release metric. Held-out task quality, calibrated decode, policy parity, red-team capacity, total cost, and upgrade regression determine whether the channel advances.
7.3 Operational superposition
Let H be a preregistered set of mutually distinguishable candidate hypotheses. A message demonstrates useful multi-hypothesis retention only if a blinded probe recovers more of the sender’s calibrated distribution over H than matched sequential and pooled baselines at the same payload and compute budget, and the receiver uses that information to improve the registered task. Dimension count, raw bytes, or a visually mixed embedding do not pass E3.
7.4 Phase coherence and cancellation
For agent i and registered hypothesis h, let the coefficient be aᵢₕ exp(jφᵢₕ). The combined message is the weighted sum of those coefficients under a fixed composition operator. Define normalized coherence:
If the denominator is zero, Cₕ is undefined: the evaluator must abstain for that hypothesis and report the zero-mass rate rather than substitute zero, one, or an epsilon. High Cₕ means phase alignment under the chosen encoding; low Cₕ means cancellation. It does not mean the hypothesis is true. E4 requires alignment to predict independently judged agreement and improve task-level consensus against matched debate and aggregation baselines. Any norm-preserving transform must satisfy U†U = I in the registered implementation.
7.5 Classical shared state
Every update has an actor, causal parent, version, permitted scope, and replayable delta. E6 measures retransmission avoided and long-horizon task effects. No non-classical correlation is asserted.
7.6 Authority invariant
A message may propose or request. Only the deterministic capability envelope and governor can authorize a side effect. Decoder confidence cannot increase authority.
Implementation Status
Status is updated for v0.4 on July 27, 2026. “Verified” below means bounded local evidence recorded under a frozen protocol; it does not mean production validation, independent replication, or a general result. The public package contains sanitized aggregate evidence, not the private product source, raw model-native state, or fixture corpus.
| Artifact | Status | Evidence | Not demonstrated |
|---|---|---|---|
| Kea fixture shell | Fixture verified | 10/10 local checks; registry self-hash; payload checks; hash-chained replay; correction; API/CLI/viewer; live fixture flag refused | Signed registry, independent decoder, external anchoring, entitlements, retention, event stream, statistical watch, learned decode |
| Waggle v0 packet | Fixture verified | 4/4 canonical packet reconstructions; content-addressed Context Pack; deterministic composition | Fair prose baseline, token reduction, accepted quality, total cost, production compatibility, covert-capacity bound |
| Byte proxy | Illustrative | 61.6% fewer bytes than a fixture prose baseline that repeats the complete Context Pack four times | E1 economics; the baseline is intentionally simple and does not use the same reference optimization |
| C14a classical phase composition | Local hardware verified | Apple M4 Max Metal composed 15 source-separated parents across six frozen numeric scenarios; CPU Kea recomputed every component; 24/24 informed outcomes matched | Quantum behavior, semantic advantage, speed advantage, learned language, production utility |
| C15d model-native context reuse | Mixed local result | One Qwen3-14B native prefix restored once and reused across six branches; Kea qualified 6/6; faster than resident full-text reconstruction from branch 3 | Cached-prefix or fresh-native advantage, overall efficiency, token/credit savings, independent replication |
| Waggle v1 | Unbuilt | Research protocol only | Training, fidelity, capacity, safety, economics |
| Learned semantic codec | Not demonstrated | Local native-state reuse is a mechanism test, not learned inter-agent language | General semantic transfer, cross-model compatibility, hidden-state deciphering |
| Live or commercial Waggle/Kea system | None | No live codec traffic, model training, paid compute, runtime deployment, or customer study | Production safety, reliability, savings, revenue, adoption |
The C15d experiment used 15 local open-weight model process loads and 690 local forward/decode steps under explicit ceilings. It used zero provider/API model calls, training runs, external calls, deployments, or authority effects. Exact measurements, the frozen completion curve, provenance hashes, and non-claims are published in RESULTS-C15D.md and data/c15d-summary.json.
Evaluation and Kill Rules
Every learned or economic claim requires paired runs on identical tasks, models, tools, authority, budgets, and stopping criteria. The prose baseline may use the same content-addressed references; Waggle does not get to claim savings that come only from a better cache.
Common protocol
- Freeze code, models, tokenizer, prompts, tools, codec, decoder, task set, rubrics, hardware, and cost conversion before evaluation.
- Use at least 30 independently instantiated tasks per workflow and increase that number if a preregistered 80% power calculation requires more.
- For learned rungs, use five seeds and report every seed. Split by task family and time; the final test set remains sealed until model and threshold selection end.
- Blind artifact judging to transport condition. Report accepted-task rate and a task-specific 0–100 quality score with inter-rater agreement.
- Report paired differences with 10,000 task-clustered bootstrap resamples and 95% intervals. Apply Holm correction across the six headline gates.
- Measure input, cached-input, and output tokens; bytes; local/GPU compute; peak memory; energy proxy; Kea overhead; latency distribution; messages; retries; loops; synthesis; and human interventions.
- Publish protocol, exclusions, failures, and negative results. No test-set rerun after a failed gate without a new version and a new untouched holdout.
| Gate | Registered question | Advance rule | Kill rule |
|---|---|---|---|
| E1 · Economics | Does v0 reduce coordination cost at accepted quality parity? | At least 30% fewer generated coordination tokens; lower 95% quality difference bound above −2 points; total cost not higher; p95 latency not higher | Quality, control, total cost, or latency misses any bound |
| E2 · Fidelity | Can Kea reconstruct or responsibly qualify each rung? | v0: 100% exact object round-trip and 100% policy test parity. Learned: preregistered critical-field macro F1 ≥0.90; selective accuracy ≥0.95 at coverage ≥0.90; independently trained decoders agree on ≥0.90 of non-abstained critical fields; ECE ≤0.05; OOD sensitivity ≥0.95 at false-positive rate ≤0.05; no authority mismatch | Any v0 mismatch, any authority/policy mismatch, or any learned accuracy, coverage, independent-agreement, calibration, or OOD bound missed |
| E3 · Multi-hypothesis capacity | Does v1 preserve useful alternatives? | At matched payload/compute, lower 95% bound for hypothesis-recall gain above zero and accepted task quality non-inferior; publish full capacity curve | No gain over sequential/pooled controls, or gain does not affect task quality |
| E4 · Phase consensus | Does registered phase composition help agreement? | Constructed agreement/contradiction AUROC ≥0.80; task quality non-inferior within 2 points; at least 20% lower total coordination cost than matched debate | Coherence fails to track agreement, or debate matches/beats total economics |
| E5 · Adversarial legibility | How much hidden content can bypass Kea? | Registered secret-recovery attack success no more than 5 points above chance while benign false-block rate ≤5%; zero authority escalation | Any authority escalation or either safety bound missed |
| E6 · Shared-state utility | Does v2 reduce retransmission without contamination? | At least 20% fewer retransmitted bytes; quality non-inferior within 2 points; 100% isolation, provenance, and policy tests | Cross-task leakage, irreplayable state, policy mismatch, or no retransmission gain |
Global stop conditions
A rung stops immediately on test leakage, unregistered model or decoder changes, missing raw evidence, unbounded payload capacity, authority mismatch, private-data contamination, inability to reproduce a result, or cost beyond the approved budget. If Kea overhead erases the gain, the transport fails. If E3 fails, “superposition” is removed from product language. If E4 fails, phase consensus is published as a negative result. If E5 fails, learned traffic does not enter production.
Four proposed workflow families
A bounded repository fix with genuinely independent leaves; a sanitized email-to-analysis-to-reviewed-draft case with sending disabled; a product-image-to-content-plan case using licensed fixtures; and a website-description-to-preview case with deployment, purchase, DNS, and commerce disabled. Each requires its own artifact rubric and data-rights record.
Threat Model
The protected assets are authority boundaries, raw message integrity, sensitive context, faithful interpretation, causal order, budgets, human corrections, and accepted artifacts. The adversary may be a malicious sender, a compromised decoder, a curious receiver, an incompatible upgrade, or an evaluator optimizing the metric rather than the task.
| Threat | Failure | Required control | Current state |
|---|---|---|---|
| Symbolic covert channel | Prose or secrets encoded in identifiers, references, numeric arrays, or content-addressed objects | Canonical vocabulary, field-level information budget, adversarial corpus, E5 | Open |
| Latent attack | Attack-associated state survives a clean-looking handoff | Independent decoder, adversarial probe, OOD watch, quarantine, fallback | Unbuilt |
| Decoder collusion | Sender and Kea agree on a benign gloss while receiver acts on hidden content | Independent decoder, behavioral probes, secret-recovery red team, disagreement | Unbuilt |
| Registry or replay tampering | Codec identity or history is rewritten coherently | User-controlled signatures, key rotation, append-only anchoring, forensic verification | Self-hash only |
| False certainty | A plausible gloss is displayed as ground truth | Calibration, alternatives, abstention, raw-link visibility, immutable correction | v0 only |
| Entitlement bypass | A decode exposes content the viewer could not access at source | Separate raw/interpretation entitlements and filtered exports | Unbuilt |
| Authority laundering | Opaque payload or gloss triggers a side effect | Deterministic capability envelope and governor; no authority in payload | Fixture invariant |
| Upgrade drift | Model, adapter, or decoder change silently changes meaning | Version pinning, regression suite, dual decode, declared fallback | IDs only |
| Shared-state contamination | One Mission changes another’s memory or leaks data | Scoped state, versioned deltas, entitlements, isolation and purge tests | Unbuilt |
Natural language is not a perfectly safe control. Steganography and ambiguity exist in prose[28]. That does not make a learned latent channel safer. The v0 hypothesis is narrower: a constrained semantic protocol may be easier to validate than unconstrained chat. The learned-rung safety hypothesis remains unproven until E5 passes.
Annex: The Quantum Analogy, Kept Honest
Quantum computing supplies useful design questions: which operations occur before readout, what information is preserved by composition, and how is measurement specified? It does not supply evidence for Waggle’s efficiency.
Categorical quantum mechanics formalizes quantum protocols with compositional structures[18]. DisCoCat uses related categorical machinery to compose distributional meanings for well-typed sentences[19]. lambeq can transform sentence-derived diagrams into tensor networks and quantum circuits[20]. These results do not imply that an arbitrary Waggle packet is “already quantum software.” A compilation milestone would require a formal grammar-to-category mapping, preservation theorem for the selected semantics, executable circuit, classical baseline, and reason the quantum representation is useful.
eQMARL studies cooperation over a genuine quantum channel and reports no explicit sharing of local observations in its architecture[21]. It is separate research, not evidence that classical shared memory behaves like entanglement. AlphaQubit and Rigetti/Quantum Machines show AI applied to quantum decoding and calibration[30][31]; they motivate decoder-first engineering only by analogy.
Roadmap and Economics
The sequence follows evidence, not spectacle. The current local experiments used no provider calls or external spend. Future hardware, data, review, and compute costs are unknown until a scoped profiler and budget are approved; this paper does not estimate them.
| Order | Deliverable | Gate | Status |
|---|---|---|---|
| 0 | Preserve baseline, claim ledger, research method, evidence, failures, and non-claims | Offline package verification and versioned public archive | v0.4 release |
| 1 | Deterministic Waggle v0 and standalone Kea replay/correction | Exact round-trip, lineage, tamper refusal, authority=false | Local fixture verified |
| 2 | Classical Apple Metal phase composition with independent CPU Kea | Real GPU receipt, matched CPU/text controls, ambiguity and cancellation refusal | C14a verified |
| 3 | Resident model-native context reuse with matched context controls | Exact quality, full startup/memory/Kea accounting, persistent break-even | C15d mixed result |
| 4 | Counterbalanced warm-start and longer-horizon native-context evaluation | Pre-output process-order freeze; beat strongest cached-prefix control after full overhead | Next falsifiable gate |
| 5 | Learned or cross-model semantic transport, if separately justified | Fidelity, capacity, adversarial legibility, matched quality and economics | Not demonstrated |
| 6 | Production or customer evaluation | Independent replication, security, reliability, rights, and explicit release authority | Not authorized by evidence |
The commercial hypothesis is that Kea could become an independently useful observability, replay, compatibility, and audit product, while a measured Waggle transport could reduce coordination overhead inside compatible systems. There is currently no customer evidence, willingness-to-pay study, production savings result, or revenue. Funding in caching and compression companies shows adjacent interest only[25][26][27].
Limitations and Non-claims
This public research release is authored by William Keenan and is not peer reviewed. The title names the research destination; the current implementation includes deterministic protocols, classical GPU phase composition, and bounded local model-native state reuse, but not a learned latent language. The landscape changes quickly; the dated v0.3 review may have missed related work. The implementation evidence has not been independently replicated, and the private source, model, native state, prompts, fixture inputs, and raw artifacts are not packaged here for public reproduction.
KE Studios does not claim:
- that language models “think” in a human-equivalent sense, or that vector dimension measures meaning;
- that Waggle is the first latent communication method, first decoder, or first use of complex representations;
- that arbitrary model families share a latent language or that closed providers expose usable internal state;
- that Kea recovers private chain-of-thought, true intent, or perfect English from learned representations;
- that the current local evidence proves token, credit, memory, energy, total-cost, safety, or commercial improvement;
- that the 61.6% byte proxy is E1, a fair protocol comparison, or a production forecast;
- that the C15d native mechanism beats the strongest cached-prefix or fresh-native control;
- that classical Metal phase composition is a learned semantic channel or establishes truth or consensus;
- that a learned, emergent, cross-model, or quantum-compilable Waggle language exists;
- quantum computation, entanglement, quantum parallelism, or quantum speedup;
- that a decoder makes latent communication safe, or that prose is inherently unsafe;
- that any message or interpretation grants authority or may trigger a consequential action;
- validated pricing, market demand, customer savings, or a compute budget.
No live Agent traffic, customer data, private message corpus, model training, provider/API calls, paid external compute, or external action was used for the local evidence described here. C15d did use bounded local open-weight inference under a frozen process/step ceiling. Author disclosure: William Keenan leads KE Studios, which may develop commercial products from Waggle and Kea. No external funding source has been reported for this release. Research content and aggregate data are licensed CC BY 4.0; the verifier code is MIT licensed. See LICENSE.
References
- Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J. & Tian, Y. Training Large Language Models to Reason in a Continuous Latent Space. COLM 2025. arXiv:2412.06769.
- Zhu, H., Hao, S., Hu, Z., Jiao, J., Russell, S. & Tian, Y. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought. NeurIPS 2025. arXiv:2505.12514.
- Rizvi-Martel, M., Rabusseau, G. & Mosbach, M. The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models. 2026 preprint. arXiv:2604.06374.
- Pham, C. et al. Let Models Speak Ciphers: Multiagent Debate through Embeddings. ICLR 2024. arXiv:2310.06272.
- Fu, T. et al. Cache-to-Cache: Direct Semantic Communication Between Large Language Models. ICLR 2026. arXiv:2510.03215v2.
- Ramesh, V. & Li, K. Communicating Activations Between Language Model Agents. ICML 2025. arXiv:2501.14082.
- Zou, J. et al. Latent Collaboration in Multi-Agent Systems. ICML 2026 Spotlight. arXiv:2511.20639v3.
- Du, Z. et al. Enabling Agents to Communicate Entirely in Latent Space. ACL 2026. arXiv:2511.09149v4.
- Liu, Y. Beyond Tokens: A Unified Framework for Latent Communication in LLM-based Multi-Agent Systems. 2026 preprint. arXiv:2606.05711v2.
- Zheng, Y., Zhao, Z. et al. Thought Communication in Multiagent Collaboration. NeurIPS 2025 Spotlight. arXiv:2510.20733.
- Wang, C., Huang, R., Sun, J., Wei, L. & Wu, Y. Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems. 2026 preprint. arXiv:2605.28214.
- Asif, S., Mohammadi Amiri, M., Abbas, M., Sattigeri, P. & Natesan Ramamurthy, K. LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems. 2026 preprint. arXiv:2605.22786.
- Miyato, T., Löwe, S., Geiger, A. & Welling, M. Artificial Kuramoto Oscillatory Neurons. ICLR 2025 Oral. arXiv:2410.13821v3.
- Plate, T. A. Holographic Reduced Representations. IEEE Transactions on Neural Networks 6(3), 1995. doi:10.1109/72.377968.
- Frady, E. P., Kleyko, D. & Sommer, F. T. Robust Computation with Rhythmic Spike Patterns. PNAS 116(36), 2019. doi:10.1073/pnas.1902653116.
- Kleyko, D. et al. A Survey on Hyperdimensional Computing aka Vector Symbolic Architectures, Part I: Models and Data Transformations. ACM Computing Surveys 55(6), 2023. doi:10.1145/3538531.
- Yıldırım, A. & Yücedağ, İ. Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks. 2026 preprint. arXiv:2512.01208v4.
- Abramsky, S. & Coecke, B. A Categorical Semantics of Quantum Protocols. LiCS 2004. arXiv:quant-ph/0402130.
- Coecke, B., Sadrzadeh, M. & Clark, S. Mathematical Foundations for a Compositional Distributional Model of Meaning. Linguistic Analysis, 2010. arXiv:1003.4394.
- Kartsaklis, D. et al. lambeq: An Efficient High-Level Python Library for Quantum NLP. 2021. arXiv:2110.04236.
- DeRieux, A. & Saad, W. eQMARL: Entangled Quantum Multi-Agent Reinforcement Learning for Distributed Cooperation over Quantum Channels. ICLR 2025. arXiv:2405.17486v2.
- Andreas, J., Dragan, A. & Klein, D. Translating Neuralese. 2017. arXiv:1704.06960v5.
- Levy, I., Paradise, O., Carmeli, B., Meir, R., Goldwasser, S. & Belinkov, Y. Unsupervised Translation of Emergent Communication. AAAI 2025. arXiv:2502.07552.
- Bronzini, M., Nicolini, C., Lepri, B., Staiano, J. & Passerini, A. Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures. 2025 preprint. arXiv:2509.25045v2.
- Tensormesh. Tensormesh Raises $20M … Bringing Total Funding to $24.5M. May 27, 2026. Official announcement.
- Y Combinator. The Token Company: Compression Middleware that Improves LLM Outputs. Winter 2026. YC company profile.
- Multiverse Computing. Multiverse Computing Raises $215M to Scale Ground-Breaking Technology that Compresses LLMs. June 12, 2025. Official announcement.
- Motwani, S. R. et al. Secret Collusion among AI Agents: Multi-Agent Deception via Steganography. 2025 revision. arXiv:2402.07510v5.
- Elhage, N. et al. Toy Models of Superposition. Anthropic, 2022. Transformer Circuits.
- Bausch, J. et al. Learning High-Accuracy Error Decoding for Quantum Processors. Nature 635, 2024. doi:10.1038/s41586-024-08148-8.
- Quantum Machines & Rigetti Computing. Quantum Machines and Rigetti Announce Successful AI-Powered Calibration of a Quantum Computer. 2026. Official announcement.
- Wei, H., Sun, Y. & Li, Y. DeepSeek-OCR: Contexts Optical Compression. 2025 preprint. arXiv:2510.18234.
- Lee, I. Y., Yang, C. & Berg-Kirkpatrick, T. Optical Context Compression Is Just (Bad) Autoencoding. 2026 revision. arXiv:2512.03643v2.