Skip to content

GLM ​

Endpoint: https://mmu.vod-qcloud.com/v1인증: VOD ApiToken / Bearer

기본 정보 ​

항목값
호출 주소https://mmu.vod-qcloud.com
인증VOD 애플리케이션에서 발급한 ApiToken / Bearer
출력텍스트 / 구조화된 응답
응답 방식JSON / SSE

지원 프로토콜 ​

규격경로대상
Chat CompletionsPOST /v1/chat/completions아래 모든 모델
MessagesPOST /v1/messages아래 모든 모델

아래 모델 ID에 해당 규격을 사용합니다. 인증과 공통 필드는 API Reference를 따릅니다.

샘플링 옵션 ​

이 문서에 나열된 모든 모델에 공통으로 적용됩니다.

파라미터설정
temperature설정 가능
top_p설정 가능

temperature는 샘플링 온도, top_p는 누적 확률 범위를 지정합니다. 한 가지 방식을 중심으로 조정합니다.

버전별 추론 설정 ​

모델 ID설정 방식비활성화
glm-5.1thinking_enabled비활성화 불가
glm-5-turbothinking_enabledfalse
glm-5thinking_enabledfalse

Chat Completions 기준입니다. 추론 내용은 reasoning_content, 최종 답변은 content에서 읽습니다. 비활성화할 수 없는 버전에는 비활성화 옵션을 전달하지 않습니다.

요청 파라미터 ​

파라미터필수타입설명
model필수String버전별 옵션 표의 모델 ID
messages필수ArrayChat Completions / Messages의 대화 입력
max_tokens선택IntegerChat Completions / Messages의 생성 한도. Messages에서는 필수
stream선택Booleantrue: SSE / false: JSON

요청 예시 ​

Chat Completions ​

json
{
  "model": "glm-5.1",
  "messages": [
    {
      "role": "user",
      "content": "Reply with exactly OK and nothing else."
    }
  ],
  "max_tokens": 512,
  "stream": false
}

응답 ​

json
{
  "id": "01a0de30-62c1-759a-a272-4c9750a9cebc",
  "object": "chat.completion",
  "model": "glm-5.1",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "OK",
        "reasoning_content": "1.  **Analyze the Request:**\n    *   Input: \"Reply with exactly OK and nothing else.\"\n    *   Constraint 1: Reply with \"OK\".\n    *   Constraint 2: Exactly that word, nothing else (no punctuation, no extra words, no formatting).\n\n2.  **Formulate the Output:**\n    *   Target string: \"OK\"\n\n3.  **Verify against constraints:**\n    *   Is it exactly \"OK\"? Yes.\n    *   Is there anything else? No.\n\n4.  **Final Output Generation:**\n    *   OK"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 13,
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    },
    "prompt_cached_tokens_details": {
      "audio_tokens": 0
    },
    "completion_tokens": 125,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "total_tokens": 138
  }
}

Messages ​

json
{
  "model": "glm-5.1",
  "messages": [
    {
      "role": "user",
      "content": "Reply with exactly OK and nothing else."
    }
  ],
  "max_tokens": 512,
  "stream": false
}

응답 ​

json
{
  "id": "01a0de33-76c0-751e-9056-bd61bd234513",
  "model": "glm-5.1",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "thinking": "1.  **Analyze the Request:** The user is asking for a specific response: \"OK\" and nothing else.\n2.  **Formulate the Output:** The output must strictly be the string \"OK\" without any additional text, punctuation (beyond what constitutes \"OK\"), or newlines if possible (though a standard single newline at the end is usually unavoidable/system-level, the raw text should just be \"OK\").\n3.  **Verify Constraints:** Exactly \"OK\". No \"Sure\", no \"OK.\", no explanation.\n4.  **Final Output Generation:** OK",
      "type": "thinking"
    },
    {
      "text": "OK",
      "type": "text"
    }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0,
    "input_tokens": 13,
    "output_tokens": 121
  }
}

특수 설정 ​

대화 이력은 messages에 순서대로 전달합니다. 추론을 반환하는 모델에서는 reasoning_content를 최종 답변과 분리합니다. 생성 한도에는 추론 토큰도 포함되므로 긴 추론 작업에는 충분한 max_tokens를 지정합니다.

基于 VitePress 构建 · 部署于腾讯云 EdgeOne Pages