Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Architect Financial Technologies announced Liquid Inference on October 7, 2026. It is an online marketplace that, according to Architect, auctions each large-language-model (LLM) inference request among providers offering the requested model. The lowest-priced eligible offer wins, subject to the buyer’s configured limits and requirements.
How does the Liquid Inference auction work?
A buyer sends a request for a model and can attach rules that determine which provider offers qualify. Architect says the request goes to providers quoting that model, and the lowest-priced offer that meets the rules receives the job. Its Auto-routing feature can select a model for a particular unit of work.
Available request controls described at launch include:
- A maximum cost per job.
- A maximum time to first token and a minimum throughput.
- Approved regions.
- A zero-data-retention requirement.
- Allow lists for providers or models.
Architect says it locks a maximum price before generation begins and charges for metered usage. Each job is also said to produce a receipt listing the winning provider, price cap, final charge, and competition depth at the time of award. These are features described in the launch materials; no independent testing of the auction, billing, receipts, or service performance is established there.
#1 Best Overall
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
How can users access it?
Architect describes Liquid Inference as available through a chat-style web interface and an API compatible with OpenAI- and Anthropic-compatible clients. The company also names official SDKs and coding tools including Claude Code, Codex, Cursor, and Cline, and says users can connect without changing their code. Compatibility and seamless integration are company claims, not independently verified results.
What was available at launch?
Architect’s October 7, 2026 launch materials reported hundreds of open- and closed-weight models. The company also reported that hundreds of tasks had completed successfully during its test phase across more than 700 models. Those are vendor-reported launch figures, not an independent benchmark or evidence of production performance.
Rank #2
- High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
- Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
- Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
- Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
- Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
The first provider cohort named by Architect comprised akashML, Boundless, engy, Grizzly, Hyperbolic, Pearl, and Zro. The launch article says providers can register models and quotes through a REST/WebSocket API, with payouts processed through Stripe. It also said providers paid no fees; provider participation and commercial terms may change.
What does it cost, and are credits available?
Architect says buyers set a per-job cost cap and pay metered usage, with the cap established before generation. The materials do not provide an independent comparison of total cost against other inference services, nor do they establish realized savings. Actual charges will depend on metered usage and the winning offer under the request’s rules.
Rank #3
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
At launch, Architect advertised $20 in credits for the first 500 users and referral credits equal to 20% of Architect fees for referred users and 10% for referrals of referrals. Its press release separately described a starting balance of inference tokens for new accounts. These are launch-era offers, so check Architect’s current terms and availability. The product introduction says credits are prepayments usable only for Architect’s service, are non-transferable, and have no cash value.
What should buyers evaluate before routing workloads?
The auction model is most useful to assess by looking at whether its controls fit a workload, rather than assuming that competitive bidding guarantees a better result. Compare the following for the requests you plan to send:
Rank #4
- 48GB AI graphics accelerator
- Cost: Set a cap that fits the job and assess the final metered charge, not just the quoted offer.
- Performance: Specify acceptable time to first token and throughput, then verify that the chosen limits suit the workload.
- Data handling and geography: Check whether zero data retention and approved-region controls meet your requirements.
- Eligibility: Limit the providers or models that may receive work when those choices matter.
- Routing: Understand whether you are naming a model or using Auto-routing to let the service choose one for a unit of work.
The launch materials document these controls but do not provide independent production comparisons or measured savings. They therefore do not establish that Liquid Inference is faster, cheaper, or higher quality than another inference option.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Availability and service limits
Architect says Liquid Inference is a service provided by Architect Financial Technologies Inc. and that independent inference providers act as subcontractors rather than contracting directly with buyers. The company says the service is not available in every jurisdiction, so prospective users should confirm local availability and applicable terms.
Best Value
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
Architect’s product article cautions that AI outputs can be inaccurate, incomplete, or inappropriate. Liquid Inference should also be distinguished from Architect’s financial-market offerings: the company describes this service as neither a financial, investment, nor digital-asset product. Other products discussed in the launch press release are subject to applicable law and regulatory approval.
Why Architect says it built the service
Architect’s stated rationale is that inference is often sold through static prices and bilateral enterprise contracts, and that suppliers on existing platforms may not have an incentive to submit their best bid in real time. That is the company’s market argument, not an independently established description of the entire inference market. Founder and CEO Brett Harrison described Liquid Inference as bringing competitive quoting, market data, and firm limit prices to individual AI requests; this is product positioning, not evidence of market-wide savings.
Sources: Architect launch press release, October 7, 2026; Architect product introduction, October 7, 2026.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

