ChatGPT
Endpoint: https://mmu.vod-qcloud.com/v1인증: VOD ApiToken / Bearer
기본 정보
| 항목 | 값 |
|---|---|
| 호출 주소 | https://mmu.vod-qcloud.com |
| 인증 | VOD 애플리케이션에서 발급한 ApiToken / Bearer |
| 출력 | 텍스트 / 구조화된 응답 |
| 응답 방식 | JSON / SSE |
지원 프로토콜
| 규격 | 경로 | 대상 |
|---|---|---|
| Chat Completions | POST /v1/chat/completions | 아래 모든 모델 |
| Responses | POST /v1/responses | 아래 모든 모델 |
| Messages | POST /v1/messages | 아래 모든 모델 |
이 문서에 나열된 모든 ChatGPT 모델은 WAND 게이트웨이의 Responses를 지원합니다. 프로토콜별 입력 필드는 API Reference를 따릅니다.
버전별 샘플링 옵션
| 모델 ID | temperature | top_p |
|---|---|---|
gpt-chat-latest | 미지원 | 미지원 |
gpt-6-sol | reasoning.effort=none일 때 설정 가능 | reasoning.effort=none일 때 설정 가능 |
gpt-6-luna | reasoning.effort=none일 때 설정 가능 | reasoning.effort=none일 때 설정 가능 |
gpt-6-astra | 미지원 | 미지원 |
gpt-5.6-terra | reasoning.effort=none일 때 설정 가능 | reasoning.effort=none일 때 설정 가능 |
gpt-5.6-sol | reasoning.effort=none일 때 설정 가능 | reasoning.effort=none일 때 설정 가능 |
gpt-5.6-luna | reasoning.effort=none일 때 설정 가능 | reasoning.effort=none일 때 설정 가능 |
gpt-5.5 | reasoning.effort=none일 때 설정 가능 | reasoning.effort=none일 때 설정 가능 |
gpt-5.4-pro | 미지원 | 미지원 |
gpt-5.4-nano | 설정 가능 | 설정 가능 |
gpt-5.4-mini | 설정 가능 | 설정 가능 |
gpt-5.4 | 설정 가능 | 설정 가능 |
gpt-5.3-codex | 설정 가능 | 설정 가능 |
gpt-5.3-chat | 미지원 | 미지원 |
gpt-5.2-chat | 미지원 | 미지원 |
gpt-5.2 | 설정 가능 | 설정 가능 |
gpt-5.1-chat | 미지원 | 미지원 |
gpt-5.1 | 설정 가능 | 설정 가능 |
gpt-5-nano | 미지원 | 미지원 |
gpt-5-mini | 미지원 | 미지원 |
gpt-5-chat | 미지원 | 미지원 |
gpt-5 | 미지원 | 미지원 |
gpt-4o | 설정 가능 | 설정 가능 |
gpt-4.1 | 설정 가능 | 설정 가능 |
temperature는 샘플링 온도, top_p는 누적 확률 범위를 지정합니다. 한 가지 방식을 중심으로 조정합니다.
버전별 추론 설정
| 모델 ID | 설정 방식 | 비활성화 |
|---|---|---|
gpt-chat-latest | 추론 설정 미지원 | — |
gpt-6-sol | reasoning.effort | none |
gpt-6-luna | reasoning.effort | none |
gpt-6-astra | reasoning.effort | 비활성화 불가 |
gpt-5.6-terra | reasoning.effort | none |
gpt-5.6-sol | reasoning.effort | none |
gpt-5.6-luna | reasoning.effort | none |
gpt-5.5 | reasoning.effort | none |
gpt-5.4-pro | reasoning.effort | 비활성화 불가 |
gpt-5.4-nano | reasoning.effort | none |
gpt-5.4-mini | reasoning.effort | none |
gpt-5.4 | reasoning.effort | none |
gpt-5.3-codex | reasoning.effort | none |
gpt-5.3-chat | 추론 설정 미지원 | — |
gpt-5.2-chat | 추론 설정 미지원 | — |
gpt-5.2 | reasoning.effort | none |
gpt-5.1-chat | 추론 설정 미지원 | — |
gpt-5.1 | reasoning.effort | none |
gpt-5-nano | reasoning.effort | 비활성화 불가 |
gpt-5-mini | reasoning.effort | 비활성화 불가 |
gpt-5-chat | 추론 설정 미지원 | — |
gpt-5 | reasoning.effort | 비활성화 불가 |
gpt-4o | 추론 설정 미지원 | — |
gpt-4.1 | 추론 설정 미지원 | — |
Responses 기준 필드는 reasoning.effort입니다. Chat Completions에서는 reasoning_effort를 사용합니다. 비활성화할 수 없는 모델은 해당 버전의 허용 추론 강도를 지정합니다.
버전별 추론 및 토큰 한도
| 모델 ID | 추론 강도 | 컨텍스트 토큰 | 최대 출력 토큰 |
|---|---|---|---|
gpt-6-sol | none / low / medium / high / xhigh / max | 1,050,000 | 128,000 |
gpt-6-luna | none / low / medium / high / xhigh / max | 1,050,000 | 128,000 |
gpt-6-astra | low / medium / high / xhigh / max | 1,050,000 | 128,000 |
gpt-5.6-terra | none / low / medium / high / xhigh / max | 1,050,000 | 128,000 |
gpt-5.6-sol | none / low / medium / high / xhigh / max | 1,050,000 | 128,000 |
gpt-5.6-luna | none / low / medium / high / xhigh / max | 1,050,000 | 128,000 |
gpt-5.5 | none / low / medium / high / xhigh | 1,050,000 | 128,000 |
gpt-5.4-pro | medium / high / xhigh | 1,050,000 | 128,000 |
gpt-5.4-nano | none / low / medium / high / xhigh | 400,000 | 128,000 |
gpt-5.4-mini | none / low / medium / high / xhigh | 400,000 | 128,000 |
gpt-5.4 | none / low / medium / high / xhigh | 1,050,000 | 128,000 |
gpt-5.3-codex | 기본 설정 | 400,000 | 128,000 |
gpt-5.2 | none / low / medium / high / xhigh | 400,000 | 128,000 |
gpt-5.1 | none / low / medium / high | 400,000 | 128,000 |
gpt-5-nano | 기본 설정 | 400,000 | 128,000 |
gpt-5-mini | 기본 설정 | 400,000 | 128,000 |
gpt-5 | minimal / low / medium / high | 400,000 | 128,000 |
gpt-4o | 기본 설정 | 128,000 | 16,384 |
gpt-4.1 | 기본 설정 | knowledge | 1,047,576 |
요청 파라미터
| 파라미터 | 필수 | 타입 | 설명 |
|---|---|---|---|
model | 필수 | String | 버전별 옵션 표의 모델 ID |
messages | 필수 | Array | Chat Completions / Messages의 대화 입력 |
input | 필수 | String / Array | Responses의 입력. messages 대신 사용 |
max_tokens | 선택 | Integer | Chat Completions / Messages의 생성 한도. Messages에서는 필수 |
max_output_tokens | 선택 | Integer | Responses의 생성 한도 |
stream | 선택 | Boolean | true: SSE / false: JSON |
요청 예시
Chat Completions
{
"model": "gpt-chat-latest",
"messages": [
{
"role": "user",
"content": "Reply with exactly OK and nothing else."
}
],
"max_tokens": 512,
"stream": false
}응답
{
"id": "01a0de31-615d-7014-9372-2ab6ab19e70e",
"object": "chat.completion",
"model": "gpt-chat-latest",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "OK"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 14,
"prompt_tokens_details": {
"audio_tokens": 0,
"cached_tokens": 0
},
"prompt_cached_tokens_details": {
"audio_tokens": 0
},
"completion_tokens": 17,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"total_tokens": 31
}
}Responses
{
"model": "gpt-6-astra",
"input": "Reply with exactly OK and nothing else.",
"max_output_tokens": 2048,
"stream": false,
"reasoning": {
"effort": "low"
}
}응답
{
"id": "resp_06ba8773e2ed38cf006ab7df36befc8195ab2906ff31076da8",
"object": "response",
"model": "gpt-6-astra",
"status": "completed",
"output": [
{
"id": "msg_06ba8773e2ed38cf006ab7df37be88819585df66bae9e722a1",
"type": "message",
"status": "completed",
"content": [
{
"type": "output_text",
"annotations": ,
"logprobs": ,
"text": "OK"
}
],
"phase": "final_answer",
"role": "assistant"
}
],
"usage": {
"input_tokens": 14,
"input_tokens_details": {
"cache_write_tokens": 0,
"cached_tokens": 0
},
"output_tokens": 5,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 19
}
}Messages
{
"model": "gpt-chat-latest",
"messages": [
{
"role": "user",
"content": "Reply with exactly OK and nothing else."
}
],
"max_tokens": 512,
"stream": false
}응답
{
"id": "01a0de34-7483-76a8-a2e6-5f997dc8388c",
"model": "gpt-chat-latest",
"type": "message",
"role": "assistant",
"content": [
{
"text": "OK",
"type": "text"
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 14,
"output_tokens": 5
}
}특수 설정
Chat Completions는 reasoning_effort, Responses는 reasoning.effort를 사용합니다. 추론 설정은 모델 버전에 맞춰 지정합니다. gpt-6-astra와 Pro 모델은 추론을 끄는 none 대신 해당 버전의 허용 강도를 사용합니다.
Responses의 output에는 추론 항목과 메시지 항목이 함께 들어갈 수 있습니다. type=message의 content에서 type=output_text를 읽습니다. Chat Completions는 choices.message.content를 읽습니다.
GPT-6 Sol의 도구 호출은 Responses 규격을 사용합니다. Chat Completions의 함수 호출은 reasoning_effort=none으로 구성합니다. 네이티브 도구와 게이트웨이 호환 필드는 구분됩니다.
버전별 추가 옵션
gpt-5.4-pro와 gpt-5.3-codex는 Responses 경로로 호출합니다. reasoning.effort로 추론 강도를 지정합니다.