All modelsOpen flagship GLM for long-horizon coding agents and million-token context work
Context 1MReasoningToolsOpen Weights
Pricing
Cache read
Reusing a prompt already cached upstream
¥2.00¥1.68/ 1M−16.0%
Cache write
Storing a prompt for later reuse
Not offered
Struck-through figures are the vendor's published list price.
The API is OpenAI-compatible — point the base URL here and pass the model id.
curl https://api.lulutokens.cn/v1/chat/completions \
-H "Authorization: Bearer $LULU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.2",
"messages": [{"role": "user", "content": "Hello"}]
}'
from openai import OpenAI
client = OpenAI(base_url="https://api.lulutokens.cn/v1", api_key="LULU_API_KEY")
response = client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Hello"}],
)
Endpoints:
openaiopenai-responseanthropic