Vision-Language Mini App
使用
Gradio
+
Transformers
展示三個常見任務:
影像描述(Image Captioning)
視覺問答(VQA)
影像-文字相似度(CLIP)
首次推理會自動下載模型;在 Spaces 可用 CPU/ZeroGPU。
1) Image Captioning
2) Visual Question Answering
3) Image-Text Similarity (CLIP)
1) Image Captioning
2) Visual Question Answering
3) Image-Text Similarity (CLIP)
上傳圖片
Drop Image Here
- or -
Click to Upload
max_new_tokens
↺
8
64
產生描述
Caption(s)
上傳圖片
Drop Image Here
- or -
Click to Upload
你的問題(英文/中文皆可)
What is in the picture?
Top-K 顯示
↺
1
5
回答問題
VQA 回答
上傳圖片
Drop Image Here
- or -
Click to Upload
候選文字:每行一個
a dog A red car A bowl of fruit people playing football
計算相似度