Hy
Endpoint: https://mmu.vod-qcloud.com/v1인증: VOD ApiToken / Bearer
기본 정보
| 항목 | 값 |
|---|---|
| 호출 주소 | https://mmu.vod-qcloud.com |
| 인증 | VOD 애플리케이션에서 발급한 ApiToken / Bearer |
| 출력 | 텍스트 / 구조화된 응답 |
| 응답 방식 | JSON / SSE |
지원 프로토콜
| 규격 | 경로 | 대상 |
|---|---|---|
| Chat Completions | POST /v1/chat/completions | 아래 모든 모델 |
| Responses | POST /v1/responses | 아래 모든 모델 |
| Messages | POST /v1/messages | 아래 모든 모델 |
아래 모델 ID에 해당 규격을 사용합니다. 인증과 공통 필드는 API Reference를 따릅니다.
샘플링 옵션
이 문서에 나열된 모든 모델에 공통으로 적용됩니다.
| 파라미터 | 설정 |
|---|---|
temperature | 설정 가능 |
top_p | 설정 가능 |
temperature는 샘플링 온도, top_p는 누적 확률 범위를 지정합니다. 한 가지 방식을 중심으로 조정합니다.
버전별 추론 설정
| 모델 ID | 설정 방식 | 비활성화 |
|---|---|---|
hy4-preview | thinking_enabled | 비활성화 불가 |
hy3 | thinking_enabled | 비활성화 불가 |
Chat Completions 기준입니다. 추론 내용은 reasoning_content, 최종 답변은 content에서 읽습니다. 비활성화할 수 없는 버전에는 비활성화 옵션을 전달하지 않습니다.
요청 파라미터
| 파라미터 | 필수 | 타입 | 설명 |
|---|---|---|---|
model | 필수 | String | 버전별 옵션 표의 모델 ID |
messages | 필수 | Array | Chat Completions / Messages의 대화 입력 |
input | 필수 | String / Array | Responses의 입력. messages 대신 사용 |
max_tokens | 선택 | Integer | Chat Completions / Messages의 생성 한도. Messages에서는 필수 |
max_output_tokens | 선택 | Integer | Responses의 생성 한도 |
stream | 선택 | Boolean | true: SSE / false: JSON |
요청 예시
Chat Completions
json
{
"model": "hy4-preview",
"messages": [
{
"role": "user",
"content": "Reply with exactly OK and nothing else."
}
],
"max_tokens": 512,
"stream": false
}응답
json
{
"id": "01a0de31-7e55-7dee-b6b5-2f5fb1918137",
"object": "chat.completion",
"model": "hy4-preview",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "OK",
"reasoning_content": "We need respond with exactly \"OK\" and nothing else. The user explicitly asks: \"Reply with exactly OK and nothing else.\" So final answer must be just OK, no extra characters, no punctuation? Exactly OK means two letters O and K. Ensure no newline? Usually can include newline? It says exactly OK and nothing else. Best final content: OK. Could be no trailing spaces. In chat, message body might have newline? To be safe, output just OK. No quotes. No markdown. Final: OK."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 29,
"prompt_tokens_details": {
"audio_tokens": 0,
"cached_tokens": 0
},
"prompt_cached_tokens_details": {
"audio_tokens": 0
},
"completion_tokens": 109,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 106,
"rejected_prediction_tokens": 0
},
"total_tokens": 138
}
}Responses
json
{
"model": "hy4-preview",
"input": "Reply with exactly OK and nothing else.",
"max_output_tokens": 512,
"stream": false
}응답
json
{
"id": "resp_a1c2722b42b84978bc7754f88a7d45a5",
"object": "response",
"model": "hy4-preview",
"status": "completed",
"output": [
{
"type": "reasoning",
"id": "rs_20260926225757nn1pd1f3",
"status": "completed",
"summary": [
{
"type": "summary_text",
"text": "We need respond exactly \"OK\" and nothing else. The user says: \"Reply with exactly OK and nothing else.\" So output only OK. No extra spaces? They said exactly OK and nothing else. Should we include newline? Usually just \"OK\". Could be trailing newline maybe okay? To be safe, output \"OK\" without newline? In chat, assistant message content likely \"OK\". We can just output OK. No reasoning."
}
]
},
{
"type": "message",
"id": "msg_20260926225757l9n1p3r5",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "OK"
}
]
}
],
"usage": {
"input_tokens": 29,
"output_tokens": 91,
"total_tokens": 120,
"output_tokens_details": {
"reasoning_tokens": 88
}
}
}Messages
json
{
"model": "hy4-preview",
"messages": [
{
"role": "user",
"content": "Reply with exactly OK and nothing else."
}
],
"max_tokens": 512,
"stream": false
}응답
json
{
"id": "01a0de34-91da-79f1-8dad-cdbedf2cd19f",
"model": "hy4-preview",
"type": "message",
"role": "assistant",
"content": [
{
"text": "OK",
"type": "text"
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 29,
"output_tokens": 111
}
}특수 설정
텍스트 대화와 코드 입력을 messages에 전달합니다. hy3와 hy4-preview는 서로 다른 모델 ID입니다. 응답에 reasoning_content가 포함되면 최종 content와 분리해서 처리합니다.