Multimodal Token Service

Usage Measured Per Modality, Reported As One Record

A single shipment file mixes documents, photographs, video clips and voice calls. Metering all of it as text tokens gives no usable view of where cost is generated. This service measures each modality on its own terms and converts everything into one comparable usage record.

Multimodal Token Service

Multimodal Token Metering & Billing

Industry workloads are rarely text-only: a single shipment file mixes documents, photographs, video clips and voice calls. This service meters each modality on its own terms and converts everything into one comparable usage record.

Text / image / video / voiceUnified conversion unitPer-modality quotasCross-modal task tracingModality-level audit log

Text Tokens

Dialogue, document summarisation and report generation metered by input and output tokens per model version.

Image Tokens

Document images, cargo condition photographs and yard stocktaking frames metered by resolution and processing depth.

Video Tokens

Handling surveillance, loading and unloading review and safety event clips metered by duration and sampling rate.

Voice Tokens

Multilingual service calls, transcription and voice broadcast metered by audio length and language pair.

Cross-Modal Task Metering

A task that reads a document, checks a photograph and drafts a reply is recorded as one traceable unit with its component costs.

Unified Conversion & Dashboard

All modalities converted to a comparable unit, with usage views by project, department, customer and modality.

Per-Modality Quota & Budget

Separate ceilings for image and video workloads, which are usually the fastest-growing cost line.

Caching & Down-Sampling

Repeated reference material served from cache, and imagery down-sampled where full resolution adds no decision value.

Modality-Level Audit

Call-level records retained per modality with model version and outcome, exportable for reconciliation.

Metering Basis

How Each Modality Is Counted

Counting rules are published rather than inferred from a bill, so a customer can estimate the cost of a workflow before it runs.

Text Tokens

Input / output tokens

Metered per model version, with prompt and completion counted separately.

Typical scenesEnquiry dialogue, document summarisation, declaration drafts, periodic reports

Image Tokens

Resolution & processing depth

Frames metered by resolution band and by whether the task is detection, classification or full field extraction.

Typical scenesDocument images, cargo condition photographs, yard stocktaking frames

Video Tokens

Duration & sampling rate

Clips metered by length and by how many frames per second are actually analysed.

Typical scenesHandling surveillance, loading and unloading review, safety event clips

Voice Tokens

Audio length & language pair

Calls and broadcasts metered by audio duration, with transcription and synthesis reported separately.

Typical scenesMultilingual service calls, transcription, voice broadcast and read-back

Governance

Control, Attribution and Audit Across Modalities

Unified Conversion Unit

Every modality is converted into one comparable unit, so a finance team reads a single usage line instead of reconciling four different measures.

Cross-Modal Task Tracing

A task that reads a document, checks a photograph and drafts a reply is recorded as one traceable unit, with component costs itemised underneath.

Per-Modality Quota & Budget

Separate ceilings for image and video workloads, which are usually the fastest-growing cost line and the easiest to overrun silently.

Caching & Down-Sampling

Repeated reference material is served from cache, and imagery is down-sampled where full resolution adds no decision value.

Modality-Level Audit Log

Call-level records retained per modality with model version, input reference and outcome, exportable for reconciliation.

Attribution by Project & Customer

Usage attributed to project, department and end customer, so recharge to the right cost centre does not require manual splitting.

Cross-Modal Workflows

Where Modalities Combine in One Task

Document Intake Desk

Image + Text

A photographed invoice is parsed into structured fields, checked against the booking record and summarised for the officer. Image and text consumption are recorded under one task ID.

Yard Condition Review

Image + Video

Still frames from gate cameras are combined with short handling clips to evidence container condition, with sampling rate reduced for routine passes.

Multilingual Service Desk

Voice + Text

An inbound call is transcribed, answered from the knowledge base and confirmed by voice broadcast, with audio length and text tokens reported together.

Safety Event Handling

Video + Text

A detected event produces a short clip, a structured record and a corrective action draft, metered as one cross-modal unit.

Talk To Us

Tell Us About Your Cargo

Share the origin, destination, cargo type and target schedule. Our team will review the details and reply with a proposed arrangement.

  • AddressLumut, Perak, Malaysia
  • TelephoneTelephone number to be confirmed
  • EmailEmail address to be confirmed
  • Working HoursMonday - Friday, 9:00 - 18:00 (MYT)

Transit time and quotation are subject to actual cargo details and carrier schedules.