Molmo 7B

MultimodaleDisponibile
di CommunityID modello: molmo-7b

Allen AI Molmo 7B-D on Replicate. Open vision-language model trained on the PixMo data, notable for pointing at and locating objects in images, not just describing them.

Prezzo
≈ 0,0505 USD/esecuzione
Input → output
Testo + Immagine → Testo
Sviluppatore
Community
Aggiornato
23 settembre 2026
01

Playground

Prova Molmo 7B

Input e output

≈ 0,0505 USD/esecuzione
Prova Molmo 7B

0 / 16.000

Impostazioni avanzate (1)
Risultato
La risposta appare qui.

Questa esecuzione

circa 0,0505 USD · 5,05 crediti

0,1513 USD (15,13 crediti) sono riservati all'inizio; il tempo GPU effettivo viene fatturato.

Per account senza acquisti precedenti: le esecuzioni oltre 2 crediti richiedono una ricarica.

Nuovo qui?

10 crediti gratuiti (0,10 USD) quando ti iscrivi con Google

Utilizzabile 24 ore dopo l'iscrizione, fino a 5 esecuzioni al giorno e al massimo 2 crediti per esecuzione. Altri metodi di accesso iniziano senza crediti.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    What do you see? Give me a detailed answer

    Response

    I see a charming scene featuring a large, fluffy white dog sitting on a wooden bench in the middle of a field. The dog appears to be a poodle mix, with curly fur covering its entire body. It's sitting upright with its front paws hanging over the edge of the bench, looking directly at the camera with a happy expression. The dog's tongue is hanging out, and its eyes are dark and expressive. The bench is made of dark brown wood with a lattice design on the backrest. It's positioned on a patch of brown, dead grass, which suggests it might be late fall or early winter. In the background, there's a vast field filled with tall, brown grass, and further back, I can see a line of bare trees against a light blue sky. The overall atmosphere is peaceful and serene. The dog seems to be enjoying its time outdoors, perhaps waiting for its owner or just taking in the scenery. The contrast between the white, fluffy dog and the brown

03

Informazioni su Molmo 7B

RiassuntoA partire da 23 settembre 2026

Molmo 7B è un modello di Community nella categoria Multimodale. Su Railwail, Molmo 7B costa ≈ 0,0505 USD per esecuzione.

Molmo 7B-D from Allen AI is an open vision-language model trained on the PixMo human-annotated dataset. Beyond captioning and visual question answering it can point to and localize specific objects in an image, returning coordinates, which is useful for grounding and agentic UI tasks. This Replicate endpoint takes an image plus a prompt and returns text answers or pointing outputs. Fully open weights and data.
04

Prezzi

Prezzi in dollari USA. L'utilizzo viene addebitato dai crediti prepagati.
Esecuzione tipica (≈ 43 s su L40S)0,0505 USD per esecuzione
Tempo GPU (L40S)0,00117 USD per secondo GPU
  • Fatturato in base al tempo GPU effettivamente utilizzato. All'avvio, 3× il prezzo tipico viene riservato dal tuo saldo e regolato successivamente.
  • 1 credito = 0,01 USD

Calcolatore di costi

Calcolatore prezzi

s

Tipico secondo il provider: circa 43,1 s

Totale

5,05 USD

505 crediti

Per esecuzione

0,0505 USD · 5,05 crediti

Fatturato in base al tempo GPU effettivo; questo è una stima.

05

API

Chiama Molmo 7B con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Nessun esempio API verificato

L'API pubblica passa un formato di input diverso da quello richiesto da questo modello. Usa il playground sopra.

06

Specifiche

ID modello
molmo-7b
Sviluppatore
Community
Categoria
Multimodale
Input
Testo, Immagine
Output
Testo
Fatturazione
In base all'utilizzo (token o tempo GPU)
Voce di catalogo aggiornata
23 settembre 2026

Parametri di input

Input e impostazioni dallo schema di input del modello. L'esempio nella sezione API mostra quali di essi l'API accetta.

  • promptObbligatorio

    Question about the image

    Tipo: Testo
    Predefinito: –
    Valori consentiti: fino a 16.000 caratteri
  • image_url

    Image URL to analyze

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • max_tokens
    Tipo: Numero intero
    Predefinito: 1024
    Valori consentiti: 1 a 4096

Etichette

  • replicate
  • allenai
  • molmo
  • vision-understanding
  • open-weights
  • grounding
07

Domande frequenti

Cos'è Molmo 7B?

Molmo 7B è un modello di Community nella categoria Multimodale.

Quanto costa Molmo 7B su Railwail?

Su Railwail, Molmo 7B costa ≈ 0,0505 USD per esecuzione. Ti viene addebitato ciò che ogni richiesta utilizza effettivamente. L'utilizzo viene pagato con crediti prepagati; 1 credito equivale a 0,01 USD.

Quali impostazioni supporta Molmo 7B?

Secondo il suo schema di input, Molmo 7B conosce questi parametri: prompt (fino a 16.000 caratteri), image_url e max_tokens (1 a 4096).

Quanto è veloce Molmo 7B?

Non ci sono ancora abbastanza esecuzioni misurate di Molmo 7B su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

Molmo 7B è migliore di BLIP?

Dipende dall'attività. Molmo 7B (Community) e BLIP (Salesforce) sono entrambi modelli nella categoria Multimodale. La pagina di confronto mostra i loro prezzi e le specifiche affiancati.

Confronta Molmo 7B e BLIP

Molmo 7B può elaborare immagini?

Sì. Molmo 7B accetta immagini come input oltre al testo.

08

Modelli comparabili

Tutti in questa categoria
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/esecuzione

    99 % più economico per unità

    Confronta Molmo 7B e BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    ≈ 0,0457 USD/esecuzione

    10 % più economico per unità

    Confronta Molmo 7B e CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    ≈ 0,0050 USD/esecuzione

    90 % più economico per unità

    Confronta Molmo 7B e Depth Anything v2

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.