๐Ÿ“„ Florence-2 Document & Image Analyzer

Upload images to analyze them with Microsoft's Florence-2 vision model.

Note: The model will be loaded automatically on first use (~5GB download, takes 2-3 minutes).

Analysis Type

Choose the type of analysis to perform

Upload an image and click 'Analyze Image' to get started!

โ„น๏ธ Ready to analyze images

โ„น๏ธ About Florence-2

Florence-2 is Microsoft's foundation vision model capable of:

  • ๐ŸŽฏ Object Detection: Identifies and locates objects with bounding boxes
  • ๐Ÿ“ Detailed Caption: Generates comprehensive descriptions of image content
  • ๐Ÿ”ค OCR: Extracts and locates text in images
  • ๐Ÿ“‹ Dense Captioning: Provides detailed captions for different regions

The model downloads automatically on first use (~5GB) and is cached for subsequent uses.

โšก Performance Notes

  • First run: Model download may take 2-3 minutes
  • GPU: Faster inference when available
  • CPU: Works but slower processing
  • Model size: ~5GB (cached after first download)
  • Supported formats: PNG, JPG, JPEG, PDF

๐Ÿ“‹ How to Use

  1. Upload a file: Click "Upload Image or PDF" and choose your file
  2. Select analysis type: Choose from the dropdown menu
  3. Click Analyze: The image will appear and you can analyze it
  4. View results: See the annotated image and detailed analysis

Good examples to try:

  • Photos with objects (cars, people, animals)
  • Screenshots with text for OCR
  • Documents or diagrams for analysis
  • Multi-object scenes for detection