ลากคลิปพูดที่อัดมาไปวาง แล้ว AI จะถอดเสียง ตัดช่วงเงียบ ใส่ซับไทย เสียงประกอบ GIF และซูมให้ จากนั้นเปิด timeline ให้คนตรวจและแก้ก่อนกดเรนเดอร์ ทั้งหมดทำงานอยู่ในเครื่องของผมเอง
Drop in a recorded talking clip and the AI transcribes it, cuts dead air, adds Thai subtitles, sound effects, GIFs and zooms, then opens a timeline for a person to review and adjust before rendering. Everything runs on my own machine.
ขั้นที่มีคนเป็นผู้ตัดสินใจหรือใช้งานSteps where a person decides or uses it


AI เขียนโค้ดให้ แต่โจทย์ เกณฑ์ และสิ่งที่ยอมไม่ได้ ผมเป็นคนกำหนด
The AI writes the code; the problem, the criteria and the non-negotiables are mine.
ผมไม่ให้ AI เรนเดอร์ออกไปเองเด็ดขาด ทุกคลิปต้องเปิด editor ให้คนตรวจก่อน เพราะงานแบรนด์โดยเฉพาะสายการเงิน ถ้าคำผิดหรือมุกไม่เหมาะหลุดออกไปครั้งเดียวก็เสียหาย
The AI never renders on its own. Every clip opens in the editor for review first, because in brand work, especially finance, one wrong word or ill-judged joke going out is costly.
คลิปแกะกล่องหรือสาธิตมักมีช่วงที่ไม่มีใครพูด แต่เป็นช่วงที่คนอยากดูที่สุด ผมเลยสั่งให้ระบบดูการเคลื่อนไหวในภาพก่อน ถ้ามีการสาธิตอยู่ให้เร่งความเร็วแทนการตัดทิ้ง
Unboxing and demo clips often have stretches with no talking that are exactly what viewers want to see. I had the system check for motion first; if something is being shown, it speeds that part up instead of cutting it.
พูดไปเลยระหว่างอัด เช่น "ซูมเข้าตรงนี้" หรือ "ใส่รูปกุ้งตรงนี้" แล้ว AI จะจับคำสั่งจากบริบทและทำให้ตรงจุด ไม่ต้องจดไว้ไปแก้ทีหลัง
Just say it while recording ("zoom in here", "put a photo here") and the AI picks up the request from context and applies it at the right moment, no notes needed.
ทุกครั้งที่ระบบเรียก AI จะปิดเครื่องมือที่งานนั้นไม่ได้ใช้ ลดจากประมาณ 30,000 token เหลือประมาณ 5,000 token ต่อการเรียกหนึ่งครั้ง และมีบันทึกการใช้ทุกงาน
Every AI call switches off tools the task doesn't need, cutting roughly 30,000 tokens down to about 5,000 per call, and usage is logged for every job.
โค้ดทั้งหมดเขียนโดย Claude Code ตามที่ผมสั่ง ผมอ่านแล้วอธิบายได้ว่าแต่ละท่อนทำอะไร และทำไมต้องเขียนแบบนี้ (ข้อมูลลับอย่างรหัสผ่านและที่อยู่เครื่องถูกตัดออก)
All code was written by Claude Code under my direction. I can explain what each part does and why it's written that way. (Secrets such as keys and machine paths are removed; code comments are in Thai.)
รวมช่วงที่มีคนพูดเข้าด้วยกัน ถ้าช่องว่างระหว่างประโยคสั้นกว่าที่ตั้งไว้ก็ไม่ตัด จุดที่สำคัญคือบรรทัดท้าย: ภาพถูกตัดทีละเฟรม แต่เสียงถูกตัดตามเวลาจริง ถ้าไม่ปัดจุดตัดให้ตรงเฟรม ความคลาดจะสะสมทุกรอยตัด จนท้ายคลิปปากไม่ตรงเสียง
Merge the spoken segments; gaps shorter than the threshold are kept. The key is the last line: video is cut frame by frame but audio by exact time, so unless every cut is snapped to a frame the drift adds up and lips fall out of sync by the end.
def compute_keep(segments, pad, min_silence, min_keep=0.25, extra=None):
"""รวม segments ที่พูด (+ ช่วงบังคับเก็บ) → ช่วงที่เก็บไว้ ตัดช่องเงียบที่ยาวกว่า min_silence"""
spans = []
for s in segments:
if not s["text"].strip():
continue
spans.append([max(0.0, s["start"] - pad), s["end"] + pad])
spans += [[a, b] for a, b in (extra or [])]
spans.sort()
merged = []
for st, en in spans:
if merged and st - merged[-1][1] < min_silence:
merged[-1][1] = max(merged[-1][1], en)
else:
merged.append([st, en])
# ปัดขอบให้ตรงเฟรม 30fps — ไม่งั้นภาพกับเสียงคลาดสะสมทุกรอยตัด
snap = lambda t: round(t * FPS) / FPS
return [(snap(s), snap(e)) for s, e in merged if snap(e) - snap(s) >= min_keep]แอปเรียก Claude Code แบบเบื้องหลังให้ช่วยอ่านบทและวางเอฟเฟกต์ แล้วให้ตอบกลับเป็น JSON ที่โปรแกรมเอาไปใช้ต่อได้ทันที ตัวเลือกที่ใส่ไว้คือการปิดเครื่องมือ ปลั๊กอิน และคำสั่งพิเศษที่ไม่จำเป็น ค่าใช้จ่ายเลยลดลงเกือบ 6 เท่า
The app calls Claude Code in the background to read the script and place effects, asking for JSON the program can use directly. The flags switch off tools, plugins and commands it doesn't need, cutting the cost almost sixfold.
def ai(prompt, read=False, model="sonnet", tries=2, turns=6):
"""ถาม claude -p → dict/list (ดึง JSON ก้อนแรกจากคำตอบ)"""
# ไม่โหลดของที่งานนี้ไม่ใช้ (เครื่องมือ/MCP/skill/settings) — ~30k tokens → ~5k
cmd = [claude_exe(), "-p", "--output-format", "json", "--model", model,
"--system-prompt", SYS, "--tools", "Read" if read else "",
"--strict-mcp-config", "--setting-sources", "project",
"--disable-slash-commands"]
...
out = json.loads(r.stdout)
log_usage(out, model) # บันทึกค่าใช้จ่ายทุกงาน
m = re.search(r"[\[{].*[\]}]", out["result"], re.S)
return json.loads(m.group(0))ซับไทยวรรณยุกต์ซ้อนกัน สระลอยผิดที่
Thai tone marks stacked or floating in the wrong place
เปิดโหมดจัดรูปอักษรแบบซับซ้อนของตัวเรนเดอร์ซับ แล้วเลือกฟอนต์ไทยที่รองรับ
Turned on complex text shaping in the subtitle renderer and chose Thai fonts that support it
ท้ายคลิปปากไม่ตรงเสียง
Lips out of sync near the end
ปัดทุกจุดตัดให้ตรงเฟรม (โค้ดด้านบน)
Snap every cut to a frame (code above)
ถอดเสียงตัดกลางคำ ซับขาดครึ่ง
Transcription split words mid-way, breaking subtitles
รวมท่อนที่ถูกตัดกลางคำกลับเข้าด้วยกันก่อนทำซับ
Re-join segments split mid-word before building subtitles
อยากใช้นอกบ้าน
Wanted to use it away from home
เปิดให้มือถือของตัวเองเข้าได้ผ่านเครือข่ายส่วนตัว (VPN) มี PIN ล็อก และมีหน้า editor ขนาดมือถือแยก
Phone access over a private network (VPN) behind a PIN, with a separate phone-sized editor