Pricing/AutoGLM-Phone-9B-Multilingual
zai-org/autoglm-phone-9b-multilingual

AutoGLM-Phone-9B-Multilingual

zai-org/autoglm-phone-9b-multilingual
Phone Agent is a mobile intelligent assistant framework built on AutoGLM, capable of understanding smartphone screens through multimodal perception and executing automated operations to complete tasks. The system controls devices via ADB (Android Debug Bridge), uses a vision-language model for screen understanding, and leverages intelligent planning to generate and execute action sequences. Users can simply describe tasks in natural language—for example, “Open Xiaohongshu and search for food recommendations.” Phone Agent will automatically parse the intent, understand the current UI, plan the next steps, and carry out the entire workflow. The system also includes: Sensitive action confirmation mechanisms Human-in-the-loop fallback for login or verification code scenarios Remote ADB debugging, allowing device connection via WiFi or network for flexible remote control and development

Features

Serverless API

Docs

zai-org/autoglm-phone-9b-multilingual is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

Available Serverless

Run queries immediately, pay only for usage

Input$0.035 / M Tokens
Output$0.138 / M Tokens

Use the following code examples to integrate with our API:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="zai-org/autoglm-phone-9b-multilingual",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=65536,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

Info

Provider
Zai-org
Quantization
-

Supported Functionality

Context Length
65536
Max Output
65536
Serverless
Supported
Input Capabilities
text, image
Output Capabilities
text