Skip to content

Grok ​

Endpoint: https://mmu.vod-qcloud.com/v1인증: VOD ApiToken / Bearer

기본 정보 ​

항목값
호출 주소https://mmu.vod-qcloud.com
인증VOD 애플리케이션에서 발급한 ApiToken / Bearer
출력텍스트 / 구조화된 응답
응답 방식JSON / SSE

지원 프로토콜 ​

규격경로대상
Chat CompletionsPOST /v1/chat/completions아래 모든 모델
ResponsesPOST /v1/responses아래 모든 모델
MessagesPOST /v1/messages아래 모든 모델

아래 모델 ID에 해당 규격을 사용합니다. 인증과 공통 필드는 API Reference를 따릅니다.

샘플링 옵션 ​

이 문서에 나열된 모든 모델에 공통으로 적용됩니다.

파라미터설정
temperature설정 가능
top_p설정 가능

temperature는 샘플링 온도, top_p는 누적 확률 범위를 지정합니다. 한 가지 방식을 중심으로 조정합니다.

버전별 추론 설정 ​

모델 ID설정 방식비활성화
gk-4.6reasoning_effort전용 비활성화 옵션 없음
gk-4.3reasoning_effort전용 비활성화 옵션 없음
gk-4.1-fast-non-reasoning비추론 모델모델 ID로 선택
gk-4-20-reasoningreasoning_effort모델 ID로 선택
gk-4-20-non-reasoning비추론 모델모델 ID로 선택
gk-4-1-fast-reasoningreasoning_effort모델 ID로 선택

Chat Completions 기준입니다. 추론 내용은 reasoning_content, 최종 답변은 content에서 읽습니다. 비활성화할 수 없는 버전에는 비활성화 옵션을 전달하지 않습니다.

요청 파라미터 ​

파라미터필수타입설명
model필수String버전별 옵션 표의 모델 ID
messages필수ArrayChat Completions / Messages의 대화 입력
input필수String / ArrayResponses의 입력. messages 대신 사용
max_tokens선택IntegerChat Completions / Messages의 생성 한도. Messages에서는 필수
max_output_tokens선택IntegerResponses의 생성 한도
stream선택Booleantrue: SSE / false: JSON

요청 예시 ​

Chat Completions ​

json
{
  "model": "gk-4.6",
  "messages": [
    {
      "role": "user",
      "content": "Reply with exactly OK and nothing else."
    }
  ],
  "max_tokens": 2048,
  "stream": false,
  "reasoning_effort": "low"
}

응답 ​

json
{
  "id": "01a0de40-a5ff-78b0-ba00-93d96f5698c7",
  "object": "chat.completion",
  "model": "grok-4.6",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "OK"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0,
      "text_tokens": 12
    },
    "prompt_cached_tokens_details": {
      "audio_tokens": 0
    },
    "completion_tokens": 1,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 114,
      "rejected_prediction_tokens": 0
    },
    "total_tokens": 127
  }
}

Responses ​

json
{
  "model": "gk-4.6",
  "input": "Reply with exactly OK and nothing else.",
  "max_output_tokens": 512,
  "stream": false
}

응답 ​

json
{
  "id": "resp_0c703ed2c27d9a19006ab7dc7cd99481949e814740a80465ed",
  "object": "response",
  "model": "grok-4.6",
  "status": "completed",
  "output": [
    {
      "id": "msg_0c703ed2c27d9a19006ab7dc80bf7481948c2eaf253c24e43e",
      "type": "message",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "annotations": ,
          "logprobs": ,
          "text": "OK"
        }
      ],
      "role": "assistant"
    }
  ],
  "usage": {
    "input_tokens": 8,
    "input_tokens_details": {
      "cache_write_tokens": 0,
      "cached_tokens": 0
    },
    "output_tokens": 116,
    "output_tokens_details": {
      "reasoning_tokens": 115
    },
    "total_tokens": 124
  }
}

Messages ​

json
{
  "model": "gk-4.6",
  "messages": [
    {
      "role": "user",
      "content": "Reply with exactly OK and nothing else."
    }
  ],
  "max_tokens": 512,
  "stream": false
}

응답 ​

json
{
  "id": "01a0de33-5952-783e-9e51-bb6ca79ca5b8",
  "model": "grok-4.6",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "text": "OK",
      "type": "text"
    }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 12,
    "output_tokens": 1
  }
}

특수 설정 ​

Chat Completions에서 reasoning_effort로 추론 강도를 지정합니다. 추론 토큰은 출력 토큰 사용량에 포함됩니다. 추론 내용을 반환하는 응답에서는 reasoning_content와 최종 content를 각각 처리합니다.

基于 VitePress 构建 · 部署于腾讯云 EdgeOne Pages