Skip to content

PixVerse ​

Endpoint: POST https://vod.intl.tencentcloudapi.comAction: CreateAigcVideoTask

기본 정보 ​

항목값
ModelNamePixVerse
ModelVersionv6 / v5.6 / c1 (소문자)
기본값ModelVersion=v6 / Resolution=1080P
가드레일 해제 지원미지원

버전별 지원 규격 ​

버전해상도비율길이
공통480P / 720P / 1080P / 2K / 4K——
c1360P / 540P / 720P / 1080P—1–15초
v6360P / 540P / 720P / 1080P—1–15초
v5.6360P / 540P / 720P / 1080P—5 / 8 / 10초; 1080P는 10초 제외

입력 조건 ​

버전조건
c1참조 / 오디오: 이미지 최대 7장, 영상 참조, 오디오 동시 생성
v6참조 / 오디오: 이미지 최대 7장, 오디오 동시 생성
v5.6참조 / 오디오: c1/v6의 참조 / 오디오 기능과 구분

요청 파라미터 ​

파라미터필수타입설명
ModelName필수String고정값 PixVerse
ModelVersion선택Stringv6 / v5.6 / c1 (소문자)
Prompt필수String생성 프롬프트
FileInfos.N선택Array참조 입력. Usage는 FirstFrame(첫 프레임) 또는 Reference(참조)
OutputConfig.Resolution선택String480P / 720P / 1080P / 2K / 4K
OutputConfig.Duration선택Integer영상 길이(초)
OutputConfig.AspectRatio선택String16:9 / 9:16 / 1:1 등

요청 예시 ​

json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "v6",
  "Prompt": "a calm sunset over the ocean, cinematic",
  "OutputConfig": {
    "Resolution": "1080P",
    "Duration": 5,
    "AspectRatio": "16:9",
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

응답 예시 ​

json
{
  "AigcVideoTask": {
    "Status": "FINISH",
    "ErrCode": 0,
    "Progress": 100,
    "Output": {
      "FileInfos": [
        {
          "FileUrl": "http://<host>.vod2.myqcloud.com/.../aigcVideoGenFile.mp4",
          "ExpireTime": "2026-08-01T10:29:48Z",
          "MetaData": {
            "Width": 1920,
            "Height": 1080,
            "Duration": 5.07,
            "Container": "mov,mp4,m4a",
            "Bitrate": 9494850
          }
        }
      ]
    }
  }
}

특수 설정 ​

Lip Sync (립싱크) ​

인물이 말하는 소스 영상에 립싱크를 입힙니다. SceneType=lip_sync로 지정하고 음성을 두 가지 모드로 지정합니다.

TTS 모드: 대사와 스피커 지정 ​

json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "lip_sync",
  "SceneType": "lip_sync",
  "Prompt": "talking",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cdn>/talking_head.mp4"
    }
  ],
  "ExtInfo": "{\"AdditionalParameters\":\"{\\\"lip_sync_tts_content\\\":\\\"안녕하세요, 반갑습니다.\\\",\\\"lip_sync_tts_speaker_id\\\":\\\"14\\\"}\"}",
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

참조 오디오 모드: 오디오 파일 입력 ​

json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "lip_sync",
  "SceneType": "lip_sync",
  "Prompt": "talking",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cdn>/talking_head.mp4"
    },
    {
      "Type": "Url",
      "Category": "Audio",
      "Url": "https://<cdn>/voice.mp3"
    }
  ],
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

소스 영상은 공개 URL 사용

소스 영상 URL은 외부에서 접근 가능한 공개 URL이어야 합니다. 프리셋 스피커는 대부분 중국어 화자 기준이라, 한국어 대사는 참조 오디오 모드를 권장합니다.

버전 선택 ​

ModelVersion특징
c1최신 버전. 캐릭터 일관성이 가장 좋고 참조 영상 입력과 립싱크를 지원
v6범용 고품질 버전
v5.6이전 세대 안정 버전. 영상 길이는 5 / 8 / 10초 중 선택

c1과 v6는 1~15초, v5.6은 5 / 8 / 10초만 지원합니다.

첫 프레임과 끝 프레임 지정 ​

FileInfos.N.Usage로 첫 프레임과 끝 프레임을 구분합니다. LastFrame을 쓰는 방식을 권장합니다.

json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "v6",
  "Prompt": "Smooth transition",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/first.jpg",
      "Usage": "FirstFrame"
    },
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/last.jpg",
      "Usage": "LastFrame"
    }
  ],
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}
Usage의미
FirstFrame첫 프레임
LastFrame끝 프레임
Reference참조 이미지

끝 프레임은 두 가지 방식이 있습니다

FileInfos.Usage=LastFrame을 쓰는 방식이 기준입니다. 최상위 LastFrameUrl도 하위 호환으로 동작하지만 새 구현에서는 Usage=LastFrame을 사용하세요. 끝 프레임 URL은 5MB 이하여야 합니다.

다중 이미지 참조 (Text / @Name) ​

참조 이미지마다 Text로 이름을 붙이고, 프롬프트에서 @Name으로 지목합니다. 이미지는 최대 7장입니다.

json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "c1",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/character.png",
      "Text": "hero",
      "Usage": "Reference"
    },
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/fan.jpg",
      "Text": "fan",
      "Usage": "Reference"
    },
    {
      "Type": "Url",
      "Category": "Image",
      "Url": "https://<cos>/earring.jpg",
      "Text": "earring",
      "Usage": "Reference"
    }
  ],
  "Prompt": "the woman in @hero slowly raises her hand and opens the @fan, while the @earring sways gently as she turns her head",
  "OutputConfig": {
    "Duration": 8,
    "AspectRatio": "3:4",
    "AudioGeneration": "Enabled",
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

@Name 뒤에는 반드시 공백

@hero runs처럼 참조 대상 이름 뒤를 띄워 사용합니다. 프롬프트의 이름과 Text 값을 일치시키고 이름은 영문 / 숫자로 작성합니다.

참조 유형 (ReferenceType) ​

Category=Video인 항목에 ReferenceType을 지정해 참조의 성격을 정합니다. GV, Kling, PixVerse에 적용됩니다.

값의미
subject피사체 중심 참조
background배경 중심 참조
json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "v5.6",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cos>/source.mp4",
      "ReferenceType": "subject"
    }
  ],
  "Prompt": "change the color of the main character's dress to white",
  "OutputConfig": {
    "StorageMode": "Permanent",
    "MediaName": "PixVerse-video-edit"
  },
  "SessionContext": "job-001"
}

립싱크 (lip_sync) ​

ModelVersion=lip_sync와 SceneType=lip_sync를 함께 지정합니다. 음성은 오디오 파일 또는 TTS로 지정합니다.

프리셋 스피커 목록 ​

ExtInfo의 lip_sync_tts_speaker_id에 아래 값 중 하나를 지정합니다. 최신 목록은 PixVerse TTS 음색 조회 API로 확인할 수 있습니다.

speaker_id설명
Auto자동 선택
2Zhen Youyu
4외국인 남성
6Li Jie
10Jiang Jianghao
11Lao Sen
12Li Jieke
13Qian Duoduo
14Wang Xiaopai
16시골 큰 목소리
18허난 사투리
19대만 억양
20산시 억양
21홍콩 억양
json
{
  "SubAppId": 123456789,
  "ModelName": "PixVerse",
  "ModelVersion": "lip_sync",
  "SceneType": "lip_sync",
  "FileInfos": [
    {
      "Type": "Url",
      "Category": "Video",
      "Url": "https://<cos>/talking_head.mp4"
    }
  ],
  "Prompt": "dancing",
  "ExtInfo": "{\"AdditionalParameters\": \"{\\\"lip_sync_tts_content\\\": \\\"lets dance and sing with me\\\", \\\"lip_sync_tts_speaker_id\\\": \\\"2\\\"}\"}",
  "OutputConfig": {
    "StorageMode": "Temporary"
  },
  "SessionContext": "job-001"
}

ExtInfo.AdditionalParameters 필드:

필드설명
lip_sync_tts_content입 모양에 맞출 대사
lip_sync_tts_speaker_id프리셋 스피커 ID

입력 규격 ​

항목제약
이미지 크기10MB 이하
이미지 포맷jpeg / jpg / png
끝 프레임 URL5MB 이하
참조 이미지최대 7장

기타 파라미터 ​

파라미터설명
NegativePrompt결과에서 제외할 요소를 설명하는 프롬프트
EnhancePrompt프롬프트 자동 보정. Enabled / Disabled
Seed난수 시드. 같은 값이면 결과를 재현할 수 있음
SessionContext콜백에 그대로 전달되는 값. 최대 1000자
SessionId중복 제거용 식별자. 3일 내 동일 ID 요청은 기존 요청 기준 처리

基于 VitePress 构建 · 部署于腾讯云 EdgeOne Pages