What Is Volantis A-1? The Photonic AI System Targeting 10,000 Tokens per Second

Volantis A-1 aims to accelerate AI inference by connecting compute and memory with light. Explore the architecture, headline targets, specifications and potential business applications.

Exploded rendering of the Volantis A-1 photonic AI inference system

What Is Volantis A-1? The Photonic AI System Targeting 10,000 Tokens per Second

What if the next leap in AI speed comes from changing how memory connects to a processor?

Volantis A-1 is a proposed AI inference system built around that question. Its approach uses light to move data between compute and memory, targeting one of the constraints behind slow AI responses: the AI memory wall.

The headline ambition is enormous. In its 1 October 2026 announcement, Volantis said A-1 is being designed for models exceeding 20 trillion parameters, with speeds of up to 10,000 tokens per second per user. These are development targets, not independently verified production results. Source: Volantis announcement

For anyone waiting while an AI coding agent works through a task, the appeal is obvious. But understanding the technology requires looking past the biggest number and asking where the waiting actually happens.

Updated: 2 October 2026.

Exploded rendering of the Volantis A-1 photonic AI inference system

Volantis A-1 system rendering. Credit: Volantis. Open full-size image.

Volantis A-1 in 30 seconds

Volantis A-1 is AI infrastructure intended to run large models quickly by improving their access to memory. It is a hardware system, rather than a new chatbot or downloadable language model.

Question Quick answer
What problem does it target? The memory bottleneck in AI inference
What is the proposed approach? Photonic interconnects between compute and memory
What makes it interesting? The ambition to increase memory capacity and bandwidth together
Is the headline speed proven? Treat it as a vendor target pending reproducible system tests

Volantis’s published architecture centres on optical links inside the accelerator. The practical question is whether those links can make a complete inference service faster and more economical. Source: Volantis overview

What is the AI memory wall?

The AI memory wall describes a mismatch between a processor’s ability to perform calculations and the memory system’s ability to supply the data it needs.

Imagine a workshop with excellent machinery, but a narrow doorway through which every component must pass. Buying another machine does little if the existing machines keep waiting for parts.

There are two separate issues:

  • Memory capacity: how much information the system can hold.
  • Memory bandwidth: how quickly that information can reach the processor.

For large language model inference, both matter. Model weights occupy memory, while the KV cache stores intermediate attention information used during generation. Longer sequences and more simultaneous requests can increase memory requirements.

NVIDIA’s own inference optimisation guide explains that token generation can be dominated by data movement. Processing a prompt, known as prefill, and generating the answer, known as decode, place different demands on hardware. Source: NVIDIA inference guide

That is why a GPU memory bottleneck can matter even when a processor has impressive computing power. The useful question is how much of that power an actual workload can use.

How does Volantis A-1 use photonics?

Volantis describes optical interconnects that extend the distance over which compute can access a fast, shared memory pool.

Its technology page contrasts electrical connections spanning approximately 2–5 millimetres with optical waveguides reaching beyond 200 millimetres. The company says this could connect more than 220 memory chiplets in one pool. That is the basis of the widely shared “220 memory chips” headline. Source: Volantis technology

The proposed benefit of this photonic memory architecture is that extra memory also contributes bandwidth. Connecting more storage would be less useful if every additional chip had to share the same narrow route to compute.

Wafer image from Volantis showing its photonic interconnect technology

Wafer image published on Volantis’s technology page. Credit: Volantis. Open full-size image.

What are VCSEL lasers?

VCSEL means vertical-cavity surface-emitting laser. Volantis uses custom micro-VCSELs as the light sources for its links, drawing on an established gallium arsenide supply chain. Its announcement describes link energy consumption below one picojoule per bit. Source: Volantis announcement

The design integrates those light sources rather than relying on external lasers. Its published link specifications include sub-five-nanosecond latency. These are component-level claims; they do not describe the total time needed to answer an AI request. Source: Volantis technology

Is this an optical computer?

A-1 combines licensed compute IP with proprietary photonics. Its photonic AI chip design focuses on data movement, rather than performing every calculation optically. Source: A-1 product page

Volantis A-1 specifications

Published A-1 system specifications:

Specification Published figure
Memory capacity 10 TB
Memory bandwidth 240 TB/s
Power envelope 20 kW
Off-wafer I/O bandwidth 10 TB/s
Rack form factor 15U

These are proposed system figures. Source: A-1 product page

Can 10 TB hold a 20-trillion-parameter model?

That depends on how parameters are represented and how the system is configured.

As a simple calculation, 20 trillion parameters at 16 bits each require approximately 40 TB for weights alone. At four bits each, the raw weight data would occupy approximately 10 TB, before metadata, cache and other runtime requirements.

This calculation is an illustration, not a description of A-1’s implementation. It explains why any demonstration should disclose quantization, model architecture, memory overhead and the number of systems used.

A parameter count, a memory figure and a speed target are not enough to reconstruct a benchmark.

What would 10,000 tokens per second actually mean?

Tokens are the pieces of text a language model processes. They are not interchangeable with words, and different tokenizers make direct comparisons imperfect. Source: NVIDIA inference guide

At a sustained 10,000 tokens per second, generating 100,000 tokens would take ten seconds. That is arithmetic, not an A-1 benchmark, and it excludes prompt processing, queueing and tool execution.

When evaluating AI inference speed, separate these measurements:

Measurement What to ask
Time to first token How long before the response begins?
Output tokens per second How quickly does generation continue?
Aggregate throughput How much output does the whole service produce across users?
End-to-end latency How long until the user’s task is finished?

For real-time AI and interactive assistants, a high system-wide throughput number may hide an uncomfortable wait for one person. For background jobs, total throughput may be the more useful measure.

An AI agent also spends time calling APIs, querying databases, running tests and waiting for external services. Faster LLM inference can shorten one part of the workflow without shortening every part equally.

Volantis A-1 benchmarks: what has been demonstrated?

Vendor claims include 15× tokens per dollar versus NVIDIA Rubin, 6× tokens per watt for low-latency 1T+ MoE workloads, and over 30× lower large-model latency. These lack independent verification here. Source: A-1 product page

The public pages reviewed for this article do not provide a reproducible, end-to-end benchmark package establishing those advantages.

A useful independent evaluation would disclose:

  • the exact model and numerical precision;
  • total parameters and active parameters per token;
  • prompt length and output length;
  • batch size and concurrent users;
  • the number of accelerators and complete system configuration;
  • measured power, software versions and comparison hardware;
  • output quality alongside speed.

For mixture-of-experts models, or MoE models, the distinction between total and active parameters is particularly relevant to interpreting a workload. Comparing headline parameter counts alone leaves too much unspecified.

Our assessment is that the architecture deserves attention, while a purchasing decision needs considerably more evidence. A compelling component demonstration and a dependable production service are different milestones.

Optical-link eye diagram published by Volantis to illustrate signal quality

Volantis’s published optical-link eye diagram. This illustrates link signal quality, not an end-to-end AI speed benchmark. Credit: Volantis. Open full-size image.

Who is behind Volantis, and how much funding has it raised?

The Volantis funding announcement covers an $88 million Series A, co-led by Lachy Groom and Abstract Ventures. Source: funding announcement

The company’s own funding post puts total capital raised at $97 million and names backers including Sam Altman, Jeff Dean, Dylan Patel and John Doerr. Source: Volantis funding post

Its leadership includes founder Tapa Ghosh and co-founder and CTO Roy Meade. The team page highlights experience in high-bandwidth memory, advanced packaging and optical communications, including Meade’s work leading Micron’s HBM programme. Source: Volantis team

That background helps explain the interest in this semiconductor startup. Funding and experience support development; customer deployments will establish whether the design delivers commercially.

Volantis A-1 release date, availability and API access

Volantis plans to deliver its first integrated inference engines to customers in 2027. The announcement does not specify a firm shipping day. Source: delivery timeline

For Volantis A-1 availability, contact the company. Its product page lists no public Volantis API endpoint. Source: A-1 access

Developers should therefore wait for documentation before assuming compatibility with an existing inference API, framework or model-serving stack. A hardware announcement is not a promise that an application can switch providers by changing one URL.

Volantis A-1 pricing: what will it cost?

Volantis A-1 pricing was not publicly listed on its product page as of 2 October 2026. Source: product information

Any future quote should be evaluated across the whole deployment. Hardware price, power, cooling, utilisation, support and software migration all affect AI infrastructure costs.

For an application owner, cost per token is only an intermediate measure. The stronger commercial measure is cost per successfully completed task.

Suppose a cheaper service needs three attempts and manual correction to complete a job. A more expensive service that finishes it correctly on the first attempt could have the better economics. That is an evaluation principle, not a claim about A-1.

Volantis A-1 vs NVIDIA GPUs: what should you compare?

A meaningful Volantis vs NVIDIA comparison needs the same workload, quality requirements and service conditions. Start with the questions below:

Comparison area Evidence to request
Model support Can the required model run without losing needed functionality?
Memory behaviour Does it sustain performance at the intended context length?
User experience What happens to latency as concurrent requests increase?
Energy efficiency What is the measured system power during useful work?
Software compatibility What must change in the deployment and monitoring stack?
Reliability How does the system behave over sustained production use?
Total cost of ownership What does a completed workload cost over the deployment’s life?

This is a more useful basis for judging an AI inference accelerator than comparing two isolated peak numbers.

What about HBM, SRAM and photonic computing?

These terms describe different parts of the problem. HBM means high-bandwidth memory. SRAM is another memory technology. Photonics describes the use of light, including the connections discussed here.

Volantis frames its work around escaping the capacity-versus-bandwidth trade-off it sees in existing memory architectures. Source: architectural explanation

A buyer should translate that ambition into a workload question: can this configuration hold the required model and working data, then serve requests at the required speed and cost?

Is Volantis a ChatGPT or Claude competitor?

It occupies a different layer. A-1 is infrastructure intended to run models. ChatGPT and Claude are familiar names for AI products and model ecosystems.

For an application developer, the relevant connection would be through supported models and deployment services. The sources reviewed here do not establish that any specific proprietary model will be offered on A-1.

Why faster inference could matter for AI coding agents

Consider a hypothetical coding agent that spends 12 minutes generating and reviewing proposed changes, plus eight minutes running tests and waiting for tools.

If inference becomes ten times faster while everything else stays the same, the job takes 9.2 minutes, down from 20. It does not finish in two minutes.

This example shows why low-latency inference could still be valuable without turning every task into an instant response. Developers could make more attempts, test more alternatives and receive feedback sooner.

It also gives a practical way to assess AI agent performance: instrument the workflow first. Measure the time spent on generation, tools, retries and review. Then identify which hardware improvements would actually remove waiting.

What could this mean for business automation and ERP?

For most businesses, the relevant opportunity is what faster AI services could enable inside operational software.

The following are proposed uses of faster inference generally. They are not announced Volantis integrations or customer case studies.

Inventory management and purchasing

An assistant could examine current stock, open sales orders, incoming purchases and supplier lead times, then prepare replenishment recommendations.

Imagine a buyer changing a supplier’s expected delivery date during a planning meeting. If inventory automation can recalculate and explain the consequences quickly, the team can compare alternatives while everyone is still present.

The practical test is whether AI inventory management improves the purchasing decision. Faster answers built on incorrect stock figures would simply produce mistakes sooner.

Manufacturing and production planning

For manufacturing automation, an assistant could investigate which customer orders are affected by a late material delivery and propose changes to the production sequence.

Useful output would identify the affected jobs, explain assumptions and show the trade-offs. Any recommendation would still need to respect material availability, machine capacity and approved delivery commitments.

Finance and exception handling

For finance automation, an assistant could examine a discrepancy between a purchase order, goods receipt and supplier invoice, then assemble the evidence for review.

The goal would be to reduce investigation time while preserving the records behind the decision. Fast text generation is valuable only if the explanation is accurate and traceable.

Across these examples, ERP integration supplies the operational context. AI workflow automation supplies analysis and permitted actions. Measuring the outcome requires tracking errors, human review time and completed work alongside response speed.

What are the main limitations to watch?

The unresolved questions concern the complete system:

  • Manufacturing: can the design be produced consistently at useful yields?
  • Integration: do compute, memory, optics and software work reliably together?
  • Deployment: what power, cooling and maintenance does installation require?
  • Performance: does the advantage persist under realistic concurrency and context lengths?
  • Economics: do savings survive full operating costs and migration effort?

These are evaluation questions, not identified defects in A-1. They describe the evidence needed to move from an interesting architecture to a dependable infrastructure choice.

What to watch next

The strongest next milestone would be a named workload running on integrated hardware, accompanied by reproducible measurements and a clear system configuration.

After that, look for customer deployments, supported-model documentation, delivery commitments and commercial terms. Independent measurements of inference latency, energy use and task cost will be more useful than another unspecific speed claim.

Frequently asked questions

What is Volantis A-1?

It is Volantis’s proposed inference system, combining compute with optical memory connections to address the AI memory bottleneck. Source: Volantis

Does Volantis already deliver 10,000 tokens per second?

That speed is presented as a target. The materials reviewed for this article do not establish an independently reproduced production result. Source: company announcement

What does “220 memory chips” refer to?

Volantis says its optical reach could connect over 220 memory chiplets in a shared pool. It is an architectural claim, not evidence that any GPU can be upgraded this way. Source: technology page

Will faster inference make AI more accurate?

Speed alone does not establish accuracy. An evaluation should hold output quality constant and check whether the faster system produces equally useful results.

Will A-1 make every AI agent task instant?

No. An agent may still need to query other systems, execute code, wait for approvals and correct mistakes. Measure the whole workflow.

Should businesses wait for photonic AI hardware before automating?

Most can assess their existing bottlenecks now. If the main delays come from disconnected records, missing information or manual handoffs, faster inference hardware would address only part of the problem.

Build the operational foundation for faster AI

Volantis is pursuing an infrastructure problem. Business owners face a related practical question: if AI could respond much faster, would it have the information and system access needed to do useful work?

A purchasing assistant needs reliable stock and supplier records. A planning assistant needs accurate production data. A finance assistant needs connected transactions and a clear approval process.

At OpsMavix, we build custom business software that connects operational workflows, including orders, stock, purchasing, production and reporting. That gives teams a practical foundation for automation as AI capabilities develop.

Explore the OpsMavix Core System to see how your operation could work in one connected system.

Sources

Getting value from OpsMavix? Add us as a preferred source on Google — you'll see more of our operations content in your AI Overviews, AI Mode and Search.