Tether Data Open-Sources VisionPsy-Nano, a 460M-Parameter On-Device Vision-Language Model
Key Takeaways
- •Tether Data has released VisionPsy-Nano, a 460-million-parameter vision-language model that processes images and text entirely on edge devices without transmitting data to remote servers.
- •The model achieved a benchmark score of 62.3, ranking first among evaluated vision-language models with fewer than 500 million parameters and outperforming competitors on 16 of 17 benchmarks.
- •Tether Data introduced a Flash variant that retained approximately 99% of the standard model's quality while generating first output tokens up to 36 times faster than selected compact models on the iPhone 15.
- •Both model versions are released under the Apache 2.0 license, allowing royalty-free use for both research and commercial applications.
- •The open-source release represents Tether's strategic expansion beyond its USDT stablecoin business into decentralized AI infrastructure built for local device deployment.

Tether Data has open-sourced VisionPsy-Nano, a 460-million-parameter vision-language model engineered to run directly on smartphones, laptops, and other edge devices without requiring continuous cloud connectivity. The release marks Tether's continued expansion beyond its core stablecoin business—USDT is the largest stablecoin by market capitalization—into open AI infrastructure.
The model was developed by QVAC, Tether Data's artificial intelligence research initiative, which focuses on building open, decentralized, and adaptive AI systems. The release reflects a broader industry effort to shift advanced multimodal capabilities from centralized data centers to devices controlled directly by users. Apple, Google, and others have similarly invested in on-device AI processing, which can reduce latency, lower server costs, and keep user data local rather than transmitting it to remote servers.
VisionPsy-Nano is designed to process and interpret both visual and textual information. Its capabilities include document analysis, optical character recognition (OCR), scene understanding, visual reasoning, and instruction following. Tether Data stated that the model delivers advanced vision-language capabilities on mobile and edge devices while operating entirely without cloud-based processing, meaning images and text processed by the model never leave the device.
Benchmark Performance
According to Tether Data's published evaluation results, VisionPsy-Nano achieved the highest overall normalized score among evaluated vision-language models with fewer than 500 million parameters. The model recorded a score of 62.3, ranking first in the sub-500-million-parameter category for overall quality and performance.
The company reported that VisionPsy-Nano performed favorably against competing compact systems on 16 of 17 benchmarks. These included models from Liquid AI and Hugging Face, as well as the base model used in VisionPsy-Nano's development.
Two Versions Address Accuracy and Speed Requirements
Tether Data released two versions of the model to support different deployment requirements.
The VisionPsy-Nano-460M version is optimized for accuracy and is intended to provide strong performance across standard vision-language evaluations at its parameter scale.
The second version, VisionPsy-Nano-460M-Flash, is designed to reduce response times for real-world mobile applications. The company said the Flash model retained approximately 99% of the full version's overall quality while achieving a normalized score of 61.4. Tether Data attributed part of the speed improvement to the use of fewer visual tokens.
In tests conducted on devices including the Google Pixel 9, Samsung Galaxy S23, Samsung Galaxy S25 Ultra, and Apple iPhone 15, the Flash version reportedly generated its first output token approximately 19 to 23 times faster than selected compact models on several Android devices. The company reported that the speed advantage reached as much as 36 times on the iPhone 15. It also said the model maintained lower latency than several competing systems tested across the selected devices.
Broad Multimodal Capabilities
Despite its relatively small size, VisionPsy-Nano ranked first across four major capability categories in the company's evaluation.
In document understanding and optical character recognition, the model was designed to identify text and structural information from complex images, including financial reports, flowcharts, and infographics. Its visual perception capabilities include scene analysis and spatial layout recognition. The company reported that the model exceeded the next-best system in its size category by a relative margin of 4.6% in that area.
Tether Data Open-Sources VisionPsy-Nano: Best-in-Class ~460M On-Device Vision-Language Model Leading Industry Benchmarks Learn more: — Tether (@tether) July 29, 2026
VisionPsy-Nano also demonstrated strong visual reasoning and knowledge performance, outperforming other models in the approximately 500-million-parameter category by a relative margin of 7.4%. The company further reported that the model performed strongly in instruction following and reliability tests. On selected evaluations, it exceeded the results of models ranging from 1.6 to 2.3 times its size.
Tether AI just released VisionPsy, a best-in-class state-of-the-art vision language model optimized for edge devices (phones, laptops, …) with the highest overall score in its class/weight — Paolo Ardoino (@paoloardoino) July 29, 2026
Tether Data said the model's benchmark results showed that compact AI systems could compete with substantially larger models while remaining suitable for direct deployment on consumer devices.
Open-Source Release and Deployment Options
Both versions have been released with open model weights under the Apache 2.0 license, one of the most widely used permissive open-source licenses, which permits both research and commercial use without royalty obligations. The company said the models are intended to support research, education, and broader developer access.
Developers can use the models through three deployment options. Full-precision inference is available through Hugging Face Transformers, while quantized GGUF versions can be deployed on mobile devices through llama.cpp. A production-oriented option using vLLM is also available for high-throughput server inference.
Tether Data has also released evaluation configurations and testing frameworks based on the VLMEvalKit protocol to support reproducibility.
Paolo Ardoino, Chief Executive Officer of Tether, said the results demonstrated that efficient, locally deployed AI could provide high-quality performance without depending entirely on centralized data centers. He indicated that making the technology openly available could give developers greater access to private AI tools operating on devices already owned by users.