Velocity • Durability

AI SOLUTIONS / CONTEXT-AWARE TIERING

Flash for GPU performance. Capacity economics for scale.

Flash is a performance medium, not a capacity medium. Placement follows file size, access pattern, and temperature with zero manual tuning: every write lands on NVMe at line rate, and cold records drain to HDD without ever going offline.

✓ Never in the write path ✓ Record-level, not file-level ✓ One namespace, no stub files
ONE VOLUME · ONE NAMESPACE · ONE SET OF BYTES
APPLICATIONS DirectFlow · NFS · SMB · S3 · Kubernetes CSI — none of them see a tier every write lands on flash, at line rate NVMe FLASH checkpoint landing zone · training reads · low-latency bursts metadata · pinned small files <256 KB · pinned hot records · eligible to drain ~90% of files are small and stay here. Nothing pinned is ever evicted. Millions of tiny files, all answering at flash speed. under flash space pressure, cold 2 MB records drain down — not whole files HDD CAPACITY bulk datasets · checkpoint history · data lake retention still online · read directly in place · no rehydration, no recall queue, no stub files The file never appears moved. Applications see a file system that costs less per TB as it grows colder.
2 MB
PLACEMENT GRANULARITY
Records drain, not whole files. A large file can be partly hot and partly cold at the same time.
256 KB
PINNED-TO-FLASH THRESHOLD
Small files stay on SSD, and metadata is never evicted at all.
25 to 30%
MORE CAPACITY PER RACK
V12 SMR HDD optimization organizes sequential zones intelligently, without compromising throughput.
20%+
COST-PER-TB REDUCTION
Delivered in V12 as an in-place upgrade for V11 customers.
THE ALL-FLASH BILL

The hyperscalers settled this argument a decade ago.

Google, Meta, and Microsoft do not run all-flash storage. They run mixed fleets: just enough NVMe to saturate GPU throughput, then draining data to high-density HDD. VDURA brings the same model to every AI factory, neocloud, and enterprise.

FLASH WASTE
Overprovisioned NVMe for data nobody reads
Buying flash for the whole dataset because part of it is hot is the single largest avoidable line item in an AI storage budget. It drains capex and inflates TCO for capacity that sits idle.
VDURA: Just enough NVMe to saturate GPU throughput, with bulk data draining to high-density HDD. Flash is sized to the working set rather than the archive.
MANUAL TUNING
Static policies cannot follow a pipeline
Access patterns shift as work moves from ingest to prep to training to inference. Hand-tuned tiering rules are stale by the next stage, and every retune is an operator task.
VDURA: Placement is decided continuously from file size, access pattern, and temperature. No manual tuning, no extra software layer, no external tiering system.
RECALL PENALTIES
Most tiers punish you for reading cold data
Stub files, rehydration steps, and recall queues turn a cold read into an outage-shaped event. Teams respond by keeping everything on flash, which defeats the point.
VDURA: Cold records are read directly from HDD in place, with no rehydration, recall queue, or stub files. Records read repeatedly warm up and move back to flash for fast access.
HOW PLACEMENT WORKS

Four rules, and none of them are yours to manage.

Policy-driven, record-level, always online. Flash for heat, HDD for bulk, and no operator in the loop between pipeline stages.

RULE 1 · INGEST
Writes land on flash
All new data hits NVMe at line rate. Tiering is never in the write path, so ingest and checkpoint bursts see pure flash speed.
RULE 2 · DRAIN
Cold records move down
Under flash space pressure, cold 2 MB records drain to HDD. More free flash means data stays hot longer.
RULE 3 · READ
Cold data is still online
Evicted records are read directly from HDD. No rehydration step, no recall queue, no stub files. Records that heat up again are promoted back to flash automatically.
RULE 4 · PIN
Some things never leave flash
Metadata is always on SSD and never evicted. Files under 256 KB are pinned to SSD, so the small-file problem that kills HDD tiers never arises.
Tier is not protocol. An S3 object is a file in the same volume, so there is no second silo and no copy job between the lake and the training tier.
Data reduction runs at the storage node, not on the client, so compression never borrows CPU or memory from your GPU nodes. Toggle it on or off at any time from the GUI or CLI.
AI-AWARE PLACEMENT AND AUTOMATION

Size, access pattern, and temperature. Continuously.

The orchestration engine reads the workload rather than a static rule table. Small and recently touched files get flash. Large sequential and infrequently touched data settles on capacity. As training moves stage to stage, placement moves with it.

Flash prioritization for metadata, small files, and recently accessed data
HDD capacity expansion for large, sequential, or infrequently accessed data
Continuous Active Capacity Balancing to eliminate hotspots and evenly distribute load
Real-time adaptation to shifting model training cycles, inference loads, and checkpoint bursts, requiring zero manual tuning
THE MEDIA LAYERS
System DRAMCACHE
Caching for unmodified data and metadata, above both flash tiers.
Intent-log SSDsINFLIGHT INTEGRITY
SSD-powered intent-log protection, replacing the legacy NVDIMMs other architectures still depend on.
Metadata SSDsNEVER EVICTED
Low-latency NVMe holding the metadata databases for rapid namespace access. VeLO lives here.
High-IOPS SSDsSMALL FILES
Small-file workloads and hot records. This is the tier that absorbs checkpoint bursts and training reads.
High-bandwidth HDDsSEQUENTIAL BULK
Large sequential reads and writes, retention, and data lake capacity. Optional expansion per node.
Commodity throughout: single-port NVMe SSDs, SATA HDDs, standard servers. No proprietary enclosure.
BLENDED ECONOMICS

The more the fleet mixes, the lower the effective rate.

Roughly 90% of files in a typical fleet are small and stay on flash, while roughly 90% of capacity is large and settles on HDD. That split is what moves the blended cost per terabyte, with no impact on the performance tier.

All flashPerformance tier rate
Maximum performance: every byte on NVMe. The right call when the whole dataset is hot.
80 / 20 blendPerformance-weighted
Flash-heavy: performance-critical workloads with early capacity relief for checkpoint history.
50 / 50 blendBalanced
Active training on flash, checkpoint history and retention on HDD.
20 / 80 blendHyperscaler model
The hyperscaler reality: hot working set on flash, bulk on HDD.
NVMe flash HDD capacity
WHAT THIS MEANS FOR THE BUSINESS

Buy flash for the working set, not for the archive.

Flash Sized To The Working Set
All-NVMe nodes for model loading, active training, and checkpointing. Flash nodes with HDD capacity expansion for archived weights, retraining inputs, and inference logs. One namespace across both.
No Second Silo
Tier is not protocol. An S3 object is a file in the same volume, so the data lake and the training tier are one capacity pool with no copy jobs between them.
Nothing To Tune
Intelligent orchestration removes the manual tiering work, the extra software layer, and the external storage system that usually comes with mixed fleets.
A Second Product To Sell
Infrequent access capacity at a fraction of the performance rate turns the same fleet into more than one line on the price sheet.
Power Goes To Accelerators
HDD carries the bulk at a fraction of the watts per terabyte. In a power-capped facility that headroom is GPU headroom.
Checkpoints Stay Cheap To Keep
Checkpoint history drains to capacity and remains online, so retaining more of it does not mean buying more flash.

See what the right flash-to-capacity ratio does to your cost per terabyte.