Vision AI logo

Vision AI

Freemium
訪問

Vision AI is Google Cloud’s set of computer-vision tools for extracting information from images, documents, and videos. It helps developers automate visual analysis with pretrained APIs, document-understanding services, video analysis, and multimodal AI options.

Vision AIとは?

Vision AI is Google Cloud’s portfolio of computer-vision and visual AI services. It helps applications interpret images, scanned documents, and videos by applying pretrained models and returning outputs such as labels, extracted text, detected objects, landmarks, faces, and content-safety signals.

The portfolio covers several related workloads. Cloud Vision API provides ready-to-use image features through REST and RPC interfaces; Document AI focuses on extracting text and structured information from documents; and Video Intelligence API analyzes stored or streaming video. Google Cloud also presents multimodal and generative options for image descriptions, classification, search, image generation, editing, and embeddings.

Use Vision AI when a product or workflow needs automated visual analysis rather than manual inspection. The right service depends on the input and desired result: basic image detection, document understanding, video analysis, or generative and multimodal processing.

Vision AIでできること

Image labeling and classification

Cloud Vision API assigns generalized labels to images and returns a description, confidence score, and topicality rating for each label.

OCR for images and documents

Text Detection recognizes text in images, while Document Text Detection is optimized for dense text, handwriting, and PDF or TIFF files. Responses can include text annotations, bounding boxes, detected languages, and document structure.

Object, face, landmark, and logo detection

Pretrained image features identify localized objects, faces, landmarks, and logos, with outputs such as coordinates, confidence scores, or bounding polygons where applicable.

Image properties and safety analysis

The API can return dominant image colors and apply SafeSearch explicit-content detection, supporting metadata and content-moderation workflows.

Document understanding

Document AI combines computer vision, natural-language processing, and machine learning to classify documents and extract text, entities, and structured data from scanned material.

Video and multimodal analysis

Video Intelligence API supports object tracking, scene understanding, activity recognition, face analysis, and text recognition in video. Other Google Cloud offerings provide visual captioning, multimodal embeddings, image generation, and image editing.

利用シーン

“Image-processing pipelines”

Add automated labels, detected objects, image properties, or safety signals to incoming images before they are indexed, reviewed, or passed to another application step.

“Scanned-document workflows”

Use document-oriented OCR and Document AI to turn PDFs, TIFF files, and document images into searchable text and structured information for downstream processing.

“Searchable video archives”

Analyze stored or streaming video for objects, scenes, activities, faces, and text so media teams can support indexing, discovery, moderation, or contextual recommendations.

“Content moderation and review”

Use SafeSearch, face and object analysis, and video-recognition capabilities to add automated signals to visual-content review processes.

“Visual descriptions and discovery”

Apply image descriptions, classification, search, or multimodal embeddings where an application needs to organize visual content or make it easier to find.

よくある質問

What is Vision AI used for?

Vision AI is used to interpret images, scanned documents, and videos. Common applications include object detection, OCR, image classification and search, content moderation, document workflows, media archives, and visual recommendations.

Which Google Cloud service should I use for image analysis?

Cloud Vision API is positioned for quick integration of common image features such as labeling, OCR, face and landmark detection, and SafeSearch. Document AI is intended for document understanding, while Video Intelligence API is intended for video analysis.

What does Cloud Vision API return for OCR?

Depending on the OCR feature, it can return recognized text, text annotations, bounding boxes, and a hierarchy covering pages, blocks, paragraphs, words, and symbols. Document Text Detection is optimized for dense text, handwriting, and PDF or TIFF files.

How is Cloud Vision API priced?

Cloud Vision API uses usage-based pricing. Each feature applied to an image is a billable unit, and the product page states that 1,000 feature units are available free each month. Google Cloud also offers broader free credits for new customers, subject to its current terms.

Can Vision AI analyze video?

Yes. Google Cloud’s Video Intelligence API analyzes stored and streaming video, including objects, scenes, activities, faces, and text. It is a separate offering from the basic image features in Cloud Vision API.

クイック情報

Category
Computer vision and visual AI
Platform
Google Cloud
Primary inputs
Images, scanned documents, PDFs, TIFF files, and stored or streaming video
Core image API
Cloud Vision API
Interfaces
Cloud Vision API is available through REST and RPC
Pricing model
Usage-based; Cloud Vision API lists 1,000 free feature units per month

Vision AIの代替品

Image Describer logo

Image Describer

imagedescriber.online

画像説明、テキスト抽出、プロンプト生成、写真の一括処理に対応するオンラインAIツール

Handwriting OCR logo

Handwriting OCR

handwritingocr.com

手書きの写真・スキャン・PDFを編集可能なテキストに変換。

Doc2X logo

Doc2X

noedgeai.com

PDFや画像をAIで処理し、編集可能な形式への変換、翻訳、API一括処理に対応

AutoCropper logo

AutoCropper

autocropper.io

スキャン画像を自動分割・補正するブラウザAIツール

LlamaParse logo

LlamaParse

www.llamaindex.ai

LlamaParse is an AI document parsing platform that converts complex PDFs, office files, spreadsheets, images, and other documents into structured, AI-ready data. It is designed for developers and enterprise teams building retrieval, extraction, and document automation workflows.

Docsumo logo

Docsumo

www.docsumo.com

Docsumo is an intelligent document processing platform for lending, insurance, healthcare, finance, and eligibility workflows. It collects, classifies, extracts, verifies, and analyzes document data, then routes exceptions and sends structured results to business systems.