Cyber threat actors have increasingly targeted AI business assets, such as credentials, cloud environments, and research, as a means of operationally exploiting AI. Now, another sophisticated method for concealing the illicit use of AI resources is coming to light. According to a report last week by security firm Team Cymru, malicious actors are using proxy servers, known as relay stations, to hide the origin of traffic to cutting-edge AI models, providing cover for model distillation attacks and potential abuse of stolen AI subscription credentials.
See also: iPhone's Stolen Device Protection deters thieves in London

Team Cymru researchers initially identified 10,867 such servers running two open-source relay platforms called Claude Relay Service (CRS) and its successor, sub2api.
Their investigation was later expanded, and the total number of such proxies is now estimated at more than 80,000. “A transfer station breaks the assumption that every edge model check is based on: that the account making a request belongs to the party consuming the response,” Scott Fisher, senior principal engineer at Team Cymru, said in the company’s report. “What we’ve uncovered is an entire ecosystem designed explicitly to break the terms and conditions of edge model providers, enabling fraud and illegal activity.” These transfer services authenticate customers with their own accounts but connect to higher-level AI services using API key pools or consumer subscriptions that in many cases have been acquired through malicious activities such as credential theft.
Team Cymru was able to link several clusters of activity through these proxies to IP addresses in China and Hong Kong. Considering that over half of the transfer gateways are hosted on VPS services in the US, it suggests that the intention is to hide the geographical location of the traffic, possibly to enable model distillation attacks, in which targeted prompts are designed to extract model knowledge that is then used to train and improve other AI models.
All of the leading labs have reported botnet attacks on their services, and earlier this month the NSA, CISA, and FBI published an advisory directly accusing six China-based AI companies of engaging in industrial-scale botnetting of US models through transfer stations and fraudulent account pools. But botnetting is far from the only use of these services.
In August, researchers from Palo Alto Networks warned of a growing number of AI token theft cases, in which stolen API tokens and subscription credentials were misused through transfer stations, leading to hundreds of thousands of dollars in losses for victims. Credentials associated with privileged enterprise developer accounts are being harvested through data theft, phishing campaigns, and supply chain attacks via npm packages.
An investigation of data dumps from information theft by Okta Threat Intelligence found 561 Anthropic session tokens collected from 5,871 infected machines, including 164 that had not expired when the dataset was released. It also identified 24 valid API keys for Gemini, OpenAI, Groq , and OpenRouter, along with underground tools designed to find AI service sessions in stolen browser data.
Meanwhile, researchers from Gambit Security observed a Chinese-speaking threat actor validating 2,975 credentials collected from 1,742 computers and uploading operational keys to an AI API resale portal. The collection included 448 Gemini keys, 254 OpenAI keys, 176 Anthropic keys, as well as credentials for Groq, OpenRouter, xAI, and Amazon Web Services.
A commercial relay ecosystem The two most common relay packages in Team Cymru's initial scan were CSR and sub2api, both published on GitHub by a developer known as Wei-Shaw.
See also: iOS 26.4: Stolen Device Protection enabled by default on all iPhones

The newer platform supports user management, per-user billing, subscription-to-API conversion, model routing, and prompt control. The sub2api project has been forked over 8,000 times, and its Telegram channel has nearly 7,000 subscribers. Interestingly, its GitHub page lists 26 commercial sponsors, including 15 relay API resellers, residential proxy vendors, two AI account providers, a content delivery network optimized for relay traffic, and a media creation API service.
Meanwhile, Palo Alto Networks found transfer stations running on new-api and one-api, two other open-source platforms, and noted that many AI transfer station ads are posted on Chinese-language marketplaces like Taobao. Team Cymru analyzed a group of servers hosted by several virtual private server providers in the US and found more than 4,000 IP addresses in China and Hong Kong connected to 304 transfer stations.
Over the course of eight days in late August, these addresses sent about 14TB to relays and received more than 7TB. The relays contacted both Chinese AI services, including DeepSeek, Qwen, Zhipu, MiniMax, and Doubao , as well as Western providers such as Anthropic, OpenAI, Google, and xAI.
However, the researchers noted that traffic to Chinese services was lower volume and heavy in downloads, while traffic to Western providers was heavy in transfers. Seventeen relays mediating traffic to Anthropic’s API uploaded about 81GB while receiving about 1.4GB, a ratio of 58 to 1. Team Cymru estimated that the uploaded data could represent between 16 billion and 23 billion quotation tokens if it consisted of text.
🔑 Secure your passwords with Proton Pass
Password manager from Proton — end-to-end encryption, passkeys, built-in 2FA, and monitoring for leaks of your credentials.
- ✔ Encrypted storage of passwords & passkeys
- ✔ Notification if any of your passwords are leaked (Dark Web Monitoring)
- ✔ Free version — on all devices
The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.
Team Cymru couldn’t see the prompts or responses, so couldn’t confirm what the sessions were about, but the unusual traffic pattern could be consistent with automated questioning for model distillation. “Prediction endpoints are live, production attack surfaces,” says Joe Brinkley, head of offensive security at Cobalt. “If providers want to protect their models from extraction, they need to defend them at the application level with real behavioral telemetry, sybil defenses, and exit checks, rather than waiting for policy and compliance to do the heavy lifting.” Protecting AI credentials Organizations should treat model credentials like other production secrets.
Long-term keys should be replaced with short-term credentials where possible, and API accounts should have configured spending limits and alerts for changes in request volume, models used, and activity at unusual times of day.
Privileged developer accounts that can create keys or change billing limits should have even tighter monitoring. Security teams should regularly scan their organization’s repositories, containers, application packages, and configuration files for exposed AI credentials. Where possible, long-lived credentials should live outside of agent sandboxes, and temporary tokens should be generated and inserted into workflows when required by approved tools.
See also: Stolen Device Protection: New security feature for iPhones

Incident response procedures should include guides for the immediate revocation and rotation of AI credentials along with the cancellation of any active session tokens associated with compromised AI subscriptions.
