Point the LLM at a bounding box.
Upload a PDF or image, say what you want found, and a vision model draws the box around it.
Your image is normalized (long edge capped at 1280px) before being sent to the model; nothing is stored beyond the detection record and the normalized copy.
Recent
View all →- pelican large.jpg · image failed
- find all circles in the image circles.png · image red circle