Build the starter library

Turn approved public product pages into a reviewable visual library.

Discover website images and rendered sections inside the exact root domain, subdomain, or product folder provided, with source provenance and approval before reuse.

  • Bounded public-site discovery
  • Images and rendered sections stay distinct
  • Content fingerprints reduce duplicates
  • Every result requires review
assets / website-discovery Review required
Discovery run

Useful visuals, not an untraceable scrape

Discovered
31
bounded attempts
New candidates
12
needs review
Unchanged
17
fingerprint match
Failed
2
reported, not hidden
Screenshot
Diagram
Brand visual

A manual discovery run is deliberately rate-limited and reports progress, additions, unchanged items, captures, and failures.

Direct answer

What is website asset discovery?

Website asset discovery is a bounded, asynchronous workflow that stays inside the root domain, subdomain, or product folder the customer provides. It prioritizes product, pricing, feature, and integration pages, then can use documentation, changelog, other first-party pages, or bounded trust pages when the useful visual library remains short. Blogged can recognize sections in semantic and modern component-based layouts, normalizes accepted visuals to WebP, records provenance and fingerprints, and places discovered items in a private review queue.

Inside Assets

A visual system with an origin, an approval state, and a deliberate use.

Each control keeps the difference between discovery, eligibility, exact placement, and generation context visible.

01
Source files

Find eligible images from the pages that explain your product.

A bounded crawl stays inside the exact website scope provided and evaluates first-party pages rather than treating the entire web as an asset source.

  • Respect a root domain, subdomain, or first-level product folder.
  • Use documentation and other first-party fallbacks only when needed.
  • Accepted JPEG, PNG, or WebP content is normalized before storage.
02
Rendered context

Capture useful marketing sections without calling them product screenshots.

Rendered public sections from semantic or div-based component layouts are stored with page type, capture kind, viewport, timestamp, and renderer provenance.

  • Homepage, pricing, feature, integration, and page sections are labeled.
  • Capture metadata distinguishes a rendered page from native application UI.
  • Overlays hidden during capture can remain recorded in provenance.
03
Change awareness

Recognize unchanged visuals and preserve the trail when a source changes.

Content hashes, source records, and last-seen timestamps keep repeat discovery useful without filling the library with silent duplicates.

  • Fingerprint matches update last-seen evidence instead of creating a copy.
  • Changed source content can point back to the asset it supersedes.
  • The run reports added, unchanged, captured, skipped, and failed results.
04
Operational control

Keep discovery bounded, site-scoped, and review-first.

Only authorized workspace roles can start discovery, every run is tied to the active site, and a cooldown prevents accidental repeat work.

  • Owner, admin, and editor roles can request a manual run.
  • A new manual run receives a fresh one-hour cooldown window.
  • Results stay private and pending until an authorized user approves them.
How it works

From a public or uploaded visual to reviewed article evidence.

The workflow preserves site scope, provenance, eligibility, and editorial control instead of turning every image into automatic generation context.

  1. 01

    Start from the active site's public URL

    Blogged validates the configured site, creates a site-scoped run, and queues the bounded discovery work.

  2. 02

    Normalize and fingerprint candidates

    Useful files and rendered sections become WebP assets with dimensions, source evidence, content hashes, and run-level progress.

  3. 03

    Review before reuse

    New candidates enter Needs review, where the team can correct metadata, approve, bulk-approve, or discard them.

A faster visual inventory without surrendering provenance or approval.

Discovery reduces manual hunting, but it deliberately stops before reuse. Public availability does not prove that a visual is useful, current, accessible, or appropriate for a specific article; the workspace review decision remains required.

FAQ

Questions about Website asset discovery.

Clear answers about discovery, approval, provenance, privacy, placement, and generation boundaries.

Does Blogged scrape every page and image on my website?

No. Discovery uses bounded page, attempt, and accepted-asset budgets. It is designed to build a useful review queue, not create an unlimited mirror of the site.

What is the difference between a website image and a website capture?

A website image is a file found in page markup. A website capture is a rendered public page section. Captures retain their page and renderer provenance and are not mislabeled as native product UI.

How does discovery avoid duplicate assets?

Blogged fingerprints normalized content and records source relationships. An unchanged fingerprint refreshes last-seen evidence instead of creating another asset, while changed source content can preserve a supersession link.

Are discovered visuals immediately available to generated posts?

No. Discovered files and captures enter a pending-review state. They become eligible for normal reuse only after an authorized workspace member approves them.

Give each article an approved visual it can rely on.

Build the library, approve what is actually useful, and choose exact placement or reference-guided generation based on the evidence the article needs.