Qwen2-VL 7B Instruct

MultimodaaliSaatavilla
kehittäjä: CommunityMallin tunnus: qwen2-vl-7b-instruct-rep

Alibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.

Hinta
≈ 0,0024 $/suoritus
Syöte → tulos
Teksti + Kuva + Video → Teksti
Kehittäjä
Community
Päivitetty
23. syyskuuta 2026
01

Leikkikenttä

Kokeile Qwen2-VL 7B Instruct

Syöte ja tulos

≈ 0,0024 $/suoritus
Kokeile Qwen2-VL 7B Instruct

0 / 16 000

Lisäasetukset (2)
Tulos
Vastaus ilmestyy tähän.

Tämä suoritus

noin 0,0024 $ · 0,24 creditiä

Alussa varataan 0,0072 $ (0,72 creditiä); todellinen GPU-aika laskutetaan.

Uusi täällä?

10 ilmaista creditiä (0,10 $) kun rekisteröidyt Googlella

Käytettävissä 24 tuntia rekisteröinnin jälkeen, enintään 5 suoritusta päivässä ja enintään 2 creditiä suoritusta kohti. Muut kirjautumismenetelmät alkavat ilman creditejä. Riittää 13 suoritukseen tästä mallista.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    Describe this video in detail.

    Response

    The video features a monkey riding a skateboard in a park. The monkey is wearing roller skates and is seen moving forward on the skateboard, with its arms outstretched for balance. The park is filled with people sitting on benches and walking around, enjoying the sunny day. The monkey appears to be enjoying itself, and the scene is filled with a sense of fun and playfulness. The park is surrounded by trees and greenery, creating a peaceful and relaxing atmosphere. Overall, the video captures a unique and entertaining moment of a monkey enjoying a day out in the park on roller skates.

  • Prompt

    Describe this in detail.

    Response

    The image shows a half of an avocado on the left side, with a pit visible inside. On the right side, there is a stylized letter "A" in a bold, uppercase font. The avocado and the letter "A" are separated by a white arrow pointing from the avocado to the letter "A," indicating a transformation or relationship between the two. The background is a light beige color.

  • Prompt

    Describe this in detail.

    Response

    The image depicts a dog, specifically a Golden Retriever, enjoying a slice of watermelon. The dog is positioned on the left side of the frame, with its head lowered towards the watermelon. The watermelon is placed on a wooden table, and the dog appears to be biting into the fruit, indicating it is eating it. The background is blurred, suggesting a natural outdoor setting with greenery. The overall scene conveys a sense of relaxation and enjoyment.

03

Tietoja: Qwen2-VL 7B Instruct

Lyhyesti23. syyskuuta 2026 alkaen

Qwen2-VL 7B Instruct on Community-kehittäjän malli kategoriasta Multimodaali. Railwailissa Qwen2-VL 7B Instruct maksaa ≈ 0,0024 $ per suoritus.

Qwen2-VL 7B Instruct is Alibaba's open-weights vision-language model packaged as a Replicate endpoint. It accepts an image (or video) plus a text prompt and answers questions, describes scenes, reads text in images and extracts structured data. Naive dynamic resolution and M-RoPE let it handle varied aspect ratios and longer visual inputs. A good self-hostable alternative to hosted VLMs for OCR and document tasks.
04

Hinnat

Hinnat Yhdysvaltain dollareissa. Käyttö laskutetaan prepaid-krediiteistä.
Tyypillinen suoritus (≈ 2 s L40S:lla)0,0024 $ per suoritus
GPU-aika (L40S)0,00117 $ per GPU-sekunti
  • Laskutus perustuu suorituksen todelliseen GPU-aikaan. Kun suoritus alkaa, 3-kertainen tyypillinen hinta varataan saldostasi ja selvitetään jälkikäteen.
  • 1 krediitti = 0,01 $

Kustannuslaskin

Hintalaskin

s

Palveluntarjoajan mukaan tyypillinen: noin 2,1 s

Yhteensä

0,24 $

24 krediittiä

Suoritusta kohti

0,0024 $ · 0,24 krediittiä

Laskutus perustuu todelliseen GPU-aikaan; tämä on arvio.

05

API

Kutsu Qwen2-VL 7B Instruct Railwail-API-avaimellasi. Käytä tätä mallin tunnusta pyynnössä:
qwen2-vl-7b-instruct-repAPI-dokumentaatioHanki API-avain

Ei vahvistettua API-esimerkkiä

Julkinen API välittää eri syötemuotoa kuin tämä malli tarvitsee. Käytä yllä olevaa leikkikenttää.

06

Tekniset tiedot

Mallin tunnus
qwen2-vl-7b-instruct-rep
Kehittäjä
Community
Kategoria
Multimodaali
Syöte
Teksti, Kuva, Video
Tuloste
Teksti
Laskutus
Käytön mukaan (tokeneja tai GPU-aikaa)
Luettelokirjaus päivitetty
23. syyskuuta 2026

Syöteparametrit

Mallin syötökaavion syötteet ja asetukset. API-osion esimerkki näyttää, mitkä niistä API hyväksyy.

  • promptpakollinen

    Question about the image or video

    Tyyppi: Teksti
    Oletus: –
    Sallitut arvot: enintään 16 000 merkkiä
  • image_url

    Image or video URL to analyze

    Tyyppi: Teksti
    Oletus: –
    Sallitut arvot: –
  • max_tokens
    Tyyppi: Kokonaisluku
    Oletus: 1024
    Sallitut arvot: 1–4 096
  • temperature
    Tyyppi: Luku
    Oletus: 0.7
    Sallitut arvot: 0–2

Tunnisteet

  • replicate
  • qwen
  • alibaba
  • vision-understanding
  • open-weights
  • ocr
07

Usein kysytyt kysymykset

Mikä on Qwen2-VL 7B Instruct?

Qwen2-VL 7B Instruct on Communityn kehittämä malli Multimodaali-kategoriassa.

Paljonko Qwen2-VL 7B Instruct maksaa Railwailissa?

Railwailissa Qwen2-VL 7B Instruct maksaa ≈ 0,0024 $ per suoritus. Sinua veloitetaan siitä, mitä kukin pyyntö todella käyttää. Käyttö maksetaan ennakkoon ostettujen krediittien avulla; 1 krediitti vastaa 0,01 $.

Mitä asetuksia Qwen2-VL 7B Instruct tukee?

Syötteen rakenteen mukaan Qwen2-VL 7B Instruct tuntee nämä parametrit: prompt (enintään 16 000 merkkiä), image_url, max_tokens (1–4 096) ja temperature (0–2).

Kuinka nopea Qwen2-VL 7B Instruct on?

Qwen2-VL 7B Instructlla ei ole vielä tarpeeksi mitattuja suorituksia Railwailissa suoritusajan ilmoittamiseksi. Se riippuu syötteestä, asetuksista ja palveluntarjoajan kuormituksesta.

Onko Qwen2-VL 7B Instruct parempi kuin BLIP?

Se riippuu tehtävästä. Qwen2-VL 7B Instruct (Community) ja BLIP (Salesforce) ovat molemmat malleja Multimodaali-kategoriassa. Vertailussa näkyvät niiden hinnat ja tekniset tiedot rinnakkain.

Vertaa Qwen2-VL 7B Instruct ja BLIP

Voiko Qwen2-VL 7B Instruct käsitellä kuvia?

Kyllä. Qwen2-VL 7B Instruct hyväksyy kuvia syötteenä tekstin lisäksi.

08

Vertailukelpoiset mallit

Kaikki tässä kategoriassa
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 $/suoritus

    88 % halvempi yksikköä kohti

    Vertaa Qwen2-VL 7B Instruct ja BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    ≈ 0,0457 $/suoritus

    1 804 % kalliimpi yksikköä kohti

    Vertaa Qwen2-VL 7B Instruct ja CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    ≈ 0,0050 $/suoritus

    108 % kalliimpi yksikköä kohti

    Vertaa Qwen2-VL 7B Instruct ja Depth Anything v2

Kaikki mallit yhden API:n kautta

Yksi API-avain kaikille Railwailin malleille. Käyttö laskutetaan prepaid-krediiteistä, 1 krediitti = 0,01 $.