multidimensional horizon
Multidimensional Vision is a one-stop AI audio and video intelligent analysis platform that converts unstructured audio and video content into structured knowledge assets. It supports local audio and video file uploads, and also supports analysis of links from mainstream platforms such as Bilibili, Douyin, and Xiaohongshu. It is aimed at individuals and teams who need to quickly extract information from audio and video.
multidimensionalview
Core parameters and statistics of multidimensional horizons
Multidimensional Vision is a one-stop AI audio and video intelligent analysis platform. Its core proposition is to convert long-form audio and video that are "inaudible and unmemorable" into searchable and reusable structured knowledge assets. The platform will be launched at the end of 2025 and will be operated by the Xiamen team to provide differentiated audio and video analysis services for C-side individual users and B-side enterprises.
| Projects | Public Information |
|---|---|
| Official positioning | AI audio and video intelligent analysis platform |
| Core value | Unstructured audio and video → Structured knowledge assets |
| Input source | Local file upload + platform link paste |
| Supported platform links | Bilibili, Douyin, Xiaohongshu, Kuaishou, Weibo, Zhihu, Youku, Xiaoyushu, Xigua, Himalaya, Tencent Conference YouTube |
| Speech recognition languages | 100+ languages, including mainstream languages and minority languages |
| Output form | Full text transcription, chapter summary, mind map, knowledge map, question and answer cards, test questions, burned subtitle video |
| Free quota | 60 minutes analysis time + 20 topic quotas |
| Membership starting price | Monthly card ¥19/month (10 hours), annual card ¥189/year (120 hours) |
| Lifetime card | ¥4999, no time limit |
| Enterprise deployment | Pure software version (containerization) and all-in-one version, support privatization |
| Usage form | Online Web platform |
| Place of Belonging | China (Fujian ICP No. 09012547-18) |
Positioning Interpretation: Audio and video information is dense but difficult to digest quickly. Multidimensional Vision is not just a "speech-to-text" tool, but it superimposes content understanding, structural reorganization and knowledge extraction capabilities on the basis of transcription, breaking down continuous audio and video into searchable, jumpable and reusable knowledge units. A 2-hour online class is distilled into structured notes within 5 minutes using AI, which is the most intuitive level of efficiency advertised.
Boundary Note: Analysis quality is affected by audio and video clarity, audio signal-to-noise ratio, and accent; dialects, noisy context, or dense professional terminology may reduce the accuracy of transcription. The free quota is only 60 minutes, and high-frequency users need to pay for extensions.
Users and market recognition of multi-dimensional vision
Gradually build user awareness in the field, and product capabilities are used by content creators and teams to improve work efficiency. Specific user scale and industry adoption data are subject to the official real-time page.
Cost advantage: Use machine analysis to replace manual repeated viewing
The cost logic of Multi-Dimensional Vision is to "replace the time of repeated manual viewing with the cost of one machine analysis." The following is a breakdown of its cost structure from three layers.
C client/individual users:
- Free File: Register to enjoy 60 minutes of analysis time + 20 topic quotas, audio and video retention for 30 days. Suitable for occasional use or physical verification.
- Monthly card ¥19/month: 10 hours of analysis time, equivalent to ¥1.9/hour. Compared to the average 3-4 hours it takes a human to watch a 2-hour video and take notes, the cost of machine analysis is negligible.
- Half-year card ¥99/half-year: 60 hours, equivalent to ¥1.65/hour. The analysis duration never expires and is not reset at the end of the month.
- Annual card ¥189/year: 120 hours, equivalent to ¥1.58/hour. Suitable for users who analyze 2-3 hours of audio and video per week.
- Lifetime Card ¥4999: Unlimited time, one step, suitable for high-frequency and heavy users.
The key design is that the analysis duration is "permanently valid, not cleared at the end of the month, and can be purchased on top of each other", which reduces users' expiration anxiety. The remaining time after the membership expires will still be retained.
Developer/API layer: Officially undisclosed API interface and pricing. In terms of product form, Multidimensional Vision currently focuses on a web platform that directly faces users and does not provide independent API services or developer SDKs. Enterprise users with API needs need to communicate through the enterprise version cooperation consultation channel, and the official reply will prevail.
Enterprise/Private Deployment:
- Pure software version: Containerized architecture, can be deployed locally on a single machine/multiple machines/clusters, and adapts to mainstream CPU/GPU environments and domestic systems.
- All-in-one version: Integrated design of software and hardware, pre-installed with optimized context and full-featured packages, ready to use out of the box, and unified maintenance. Supports concurrent analysis of up to 30 channels of video, and can complete in-depth analysis of a single video within 1 minute at the fastest, with 7×24-hour business continuity.
- The enterprise version has built-in capabilities such as speech recognition, face recognition, image OCR, summary summary, deep forgery detection, etc. It also supports custom algorithms, models and LOGOs.
- Pricing is subject to business communication, and standardized prices have not been disclosed.
Hidden costs: Secondary proofreading time when the transcription accuracy is insufficient, the cost of manual review of complex audio, and the risk of membership time consumption not matching actual demand. It is recommended to use the free quota to test the transcription quality of typical scenarios first, and then decide on a paid plan after confirming the availability.
Main functions of Multidimensional Vision
Multidimensional Vision's capability system covers the entire link of "audio and video input → AI analysis → structured output → content creation", and there is a clear synergy between each function.
-
High-precision speech transcription: Convert the speech content in audio and video into editable text, supporting 100+ languages. The synergistic value is that the transcribed text is the basic raw material for all subsequent analysis (summaries, mind maps, Q&A), and the accuracy of transcribing directly determines the downstream quality.
-
Multi-language translation: Provides cross-language translation based on transliteration, covering mainstream languages such as Chinese, English, Japanese, and Korean, as well as small languages in Southeast Asia, Northern Europe, and Africa. Synergy effect: A foreign language video can simultaneously produce original text transcription + Chinese translation + bilingual subtitles, solving the multi-step switching between "understanding foreign language content" and "producing Chinese notes".
-
Intelligent summary and chapter overview: AI automatically extracts the core points of the full text and generates chapter segments based on content themes. Synergy effect: The chapter overview is linked with the mind map. Users can directly jump to the corresponding video location through the mind map, achieving "on-demand viewing" instead of linear playback.
-
Visual analysis: including OCR recognition (titles, subtitles, channel LOGO, advertising slogans in the video screen), face analysis, and image analysis. Synergy effect: When analyzing teaching videos or product launches, OCR can extract text information in PPT and complement each other with voice transcription, forming a dual-channel information extraction of "screen text + voice content".
-
AI Dialogue and Knowledge Questions and Answers: Conduct multiple rounds of in-depth questions and answers based on audio and video analysis content. Users can ask for details, ask for examples or comparisons. Synergy: This feature extends the results of a single analysis into an interactive knowledge base rather than a one-time output.
-
Interactive Learning Tools (Flashcards & Quiz): AI automatically generates question and answer cards and test questions from the core knowledge points of the course. Synergy effect: From "watching the video" to "consolidating memory" is completed on the same platform, suitable for self-assessment and review in educational scenarios.
-
Video subtitle burning: Squeeze the AI-generated subtitles (after translation) into the video screen with one click to generate a new video file with subtitles. Implicit value: This function extends the analysis results from "text notes" to "distributable videos", which is particularly critical for content creators' foreign language transfer scenarios.
-
Customized templates and prompt words: Supports preset analysis templates (course study, character interviews, content creation, meeting minutes, etc.), as well as custom prompt words to customize AI analysis dimensions. Synergy: The combination of template + custom prompt words allows the same audio and video to produce multiple analysis reports from different perspectives, adapting to the consumer needs of different positions.
-
Speaker identification and to-do extraction: Automatically distinguish different speakers and extract to-do items from meeting recordings. Synergy effect: The combination of speaker identification + to-do extraction + summary enables the output of meeting minutes to achieve "minutes are produced immediately after the meeting".
Model and version evolution of multi-dimensional vision
Multidimensional Vision operates as an online SaaS platform and does not have a public version number system. The following records its evolution according to public milestones.
Platform launch period (~November 2025)
The product provides external services for the first time, and its core capabilities focus on audio and video uploading, platform link analysis, speech transcription and structured summary output. In the initial stage, the input sources are mainly Bilibili, Douyin, and Xiaohongshu.
Function expansion period (~first half of 2026)
Advanced functions such as visual analysis (OCR, face, image analysis), AI dialogue Flashcards/Quiz interactive learning, video subtitle burning, custom templates and prompt words will be launched in batches. Link coverage has been extended to 12 platforms including Kuaishou, Weibo, Zhihu, Youku, Xiaoshi Universe, Xigua Video, Himalaya, and Tencent Conference YouTube.
Enterprise version released (~2026)
Launched two enterprise deployment solutions, pure software version and all-in-one version, supporting privatized deployment, localized adaptation of 30-channel concurrent analysis, built-in deep forgery detection and other capabilities. The release of the enterprise version marks the extension of the product from a pure C-side tool to B-side scenarios.
Version context summary: The version evolution direction of Multi-dimensional Vision is "input source expansion → analysis depth enhancement → output form diversification → enterprise deployment method". The official unified version number and precise date have not been disclosed. The above milestones are organized according to the information on the public page.
Technical advantages of multi-dimensional vision
The technical advantages of Multi-Dimensional Vision are based on the architecture of "multi-modal understanding + analysis link closure", rather than a single algorithm indicator.
Mechanism: The platform connects the four model capabilities of speech recognition (ASR), natural language processing (NLP), machine translation (MT), and computer vision (CV) into an analysis pipeline. The audio and video go through multiple stages of "sound track separation → speech transcription → paragraph segmentation → summary generation → knowledge point extraction → structured assembly", and the output of each stage becomes the input of the next stage.
Effect:
- Integrated Transcription + Understanding: In the traditional process, users need to first use one tool to convert speech to text, and then copy it to another tool for summary and notes. Multidimensional Vision compresses this chain into a single upload/paste operation, eliminating cross-tool switching loss.
- Visual + Voice Dual Channel: OCR extraction of screen text and ASR extraction of voice content complement each other. For example, when analyzing a press conference video, the chart titles in the PPT are captured through OCR, the speaker's explanation is transcribed through ASR, and the two are integrated and presented in the final analysis report.
- 100+ language coverage: The platform claims to support the recognition and translation of 100+ languages around the world, including small languages in Southeast Asia, Africa, Northern Europe and other regions. This coverage is of practical value to creators researching cross-border content or targeting overseas users.
- Enterprise-level concurrency: The all-in-one solution supports concurrent analysis of 30 channels of video. A single video can complete in-depth analysis within 1 minute at the fastest, and has 7×24 hours of uninterrupted operation capabilities.
Applicable scenarios: The mechanism determines that the platform is most suitable for scenarios with "diversified input sources, structured output requirements, and multiple rounds of retrieval", rather than simple real-time subtitles or live translation.
Price and Limitations:
- Transcription accuracy is greatly affected by audio quality - background noise, overlapping dialogue, and non-standard accents may lead to transcription errors, which will amplify the analysis chain to the point where summary and knowledge extraction are inconsistent.
- Video length is linked to membership time. A single analysis of long videos (such as courses or meetings exceeding 2 hours) consumes more time quota.
- The specific model names and versions used by the platform are not disclosed, and users cannot independently evaluate its technical roadmap and benchmark performance.
How to use multi-dimensional vision
Multidimensional Vision currently only provides complete analysis functions through the web page, without the need to install a client.
| Entrance | Access method | Prerequisites |
|---|---|---|
| Use the official website directly | Open https://dwsj.cn/ in the browser, register/log in | Valid email or mobile phone number |
| Link analysis | Paste the target platform video link on the analysis page | The link must be publicly accessible |
| Local file upload | Select local audio and video files to upload on the analysis page | File format and size are limited by the platform |
| Enterprise version application | Submit cooperation consultation through the official website enterprise version page | Business communication is required to confirm requirements |
Typical usage process:
- Register an account and log in to the Multidimensional Vision Web platform.
- Select the input method: paste the platform video link, or upload local audio and video files.
- Wait for AI automatic analysis (the time depends on the length of the audio and video and the server load).
- Check the analysis results: full text transcription, chapter summary, mind map, keyword extraction, speaker identification, etc.
- Advanced operations: Use AI dialogue to ask for details, generate Flashcards/Quiz self-tests, apply custom templates to adjust the output format, and use the subtitle burning function to export videos with subtitles if you are processing foreign language videos.
- Export results: copy text, export mind map, download video with burned subtitles.
Recommendations for first-time users: First use the free quota (60 minutes) to test 1-2 clear audio and video segments to verify the transcription accuracy and structured output quality, and then decide whether to purchase the membership extension. Link analysis gives priority to mature platforms such as Bilibili and YouTube, which have more stable compatibility.
Multi-dimensional vision product pricing
Multidimensional Vision's pricing system is divided into two sets of logic: individual membership and corporate deployment.
Individual membership pricing
| Package | Price | Equivalent to monthly price | Including analysis time | Features |
|---|---|---|---|---|
| Free trial | ¥0 | — | 60 minutes + 20 topics | Audio and video retention for 30 days, suitable for quality verification |
| Monthly card | ¥19/month (original price ¥39) | ¥19 | 10 hours | Monthly as needed, duration never expires |
| Half-year card | ¥99 (original price ¥199) | ¥16.5 | 60 hours | Duration can be stacked, suitable for moderate users |
| Annual card | ¥189 (original price ¥389) | ¥15.75 | 120 hours | The most cost-effective, suitable for high-frequency users |
| Lifetime card | ¥4999 | — | No time limit | One step, suitable for heavy users/teams |
Billing Mechanism Description:
- The "analysis duration" of all packages is permanently valid and will not be cleared at the end of the month. The remaining duration will be retained after the membership expires.
- Packages can be purchased in combination, and the duration and validity period are automatically added to the existing ones.
- Invite friends to get free time rewards.
- Processing priority: Free < Monthly Card/Semi-Annual Card < Annual Card < Lifetime Card (membership tasks will be processed first during peak hours).
Enterprise Deployment Pricing
The enterprise version provides two solutions: pure software version and all-in-one version, both of which adopt business communication pricing. Compared with the personal version, the feature set adds enterprise-level features such as custom algorithms and models, deep forgery detection, localized system adaptation, and 30-way concurrent support. The specific price needs to be obtained through the official corporate page or the cooperation consultation channel. There is no public standard quotation.
API Pricing: Officially undisclosed API services and billing standards. Currently, the product does not export analysis capabilities in the form of API.
Application scenarios of multi-dimensional vision
The ability of multi-dimensional vision produces differentiated value in different positions and tasks. The following is a summary of the "tasks + benefits + verification points" of four typical scenarios.
-
Content secondary creation and multi-platform distribution: Full-time bloggers or new media operators import Bilibili/YouTube video links into multi-dimensional vision, use Xiaohongshu note templates or custom templates to automatically generate copywriting with emoticons and hashtags, or use the subtitle burning function to generate Chinese subtitles versions for foreign videos. Key points to check: The usability of the template output copy - whether it still requires a lot of manual rewriting. Judging from user testimonies, the content creation time for a 10-minute video can be shortened from 1 hour to 3 minutes, but the consistency of the copywriting style and brand voice still requires manual control.
-
Learning and Course Review: Students or self-learners convert online class recordings (local upload or platform link) into structured notes, and use Flashcards and Quiz functions to generate self-test questions for review and consolidation. Key points of verification: The correlation between the accuracy of transcription of professional terms and the extraction of knowledge points. For courses involving dense terminology (such as medicine, law), manual review of key definitions and concepts is required for accuracy.
-
Meeting minutes and interview organization: Product managers and researchers import meeting recordings or expert interview recordings, and use speaker identification, summary generation and to-do extraction functions to quickly produce structured minutes. Verification focus: The spokesperson distinguished between accuracy and completeness of to-do extraction. The organizing time for a 2-hour weekly meeting can be reduced from 1 hour to 5 minutes, but for complex meetings involving multi-party games, the AI summary may miss implicit disagreements or unstated conclusions.
-
Industry Research and Intelligence Analysis: Securities researchers and analysts use customized prompt word templates to extract dimensions such as "market views", "risk reminders" and "data predictions" from a large number of industry conference calls to form structured research report materials. Key points of verification: The effect of the prompt word template in guiding the analysis perspective - the difference and controllability of the output results under different prompt word settings.
-
HR interview review: HR uses speaker summary, mind mapping and emotion analysis functions to conduct a structured review of the group interview video to refine the key points and logical lines of each candidate's speech. Verification Points: Objectivity and privacy compliance boundaries of sentiment analysis - In scenarios involving candidate evaluation, AI analysis results are only used as a reference and not as a basis for decision-making.
Applicable people for multi-dimensional vision
-
Content Creators and New Media Operations: Individual creators and operations teams who need to frequently convert long videos into graphic notes, distribute them across multiple platforms, or process foreign language video materials. The product provides significant efficiency improvements in the two directions of "video → copywriting" and "foreign language → Chinese". Not suitable for boundaries: For high-quality content that has high requirements for original visual design, the AI-generated templated copy may lack brand uniqueness and requires manual in-depth customization.
-
Students and lifelong learners: Online course learners, postgraduate entrance examination/certificate exam preparers need to extract key points from a large number of course videos and test themselves. Flashcards and Quiz functions create differentiated value in the "Understand → Consolidate" relationship. Unsuitable Boundary: For humanities courses that require in-depth thinking or critical thinking, AI summaries may simplify complex arguments and are recommended as a preview/review aid rather than a replacement for intensive reading.
-
Corporate professionals (product managers, HR researchers): workplace professionals whose daily output involves audio and video analysis tasks such as meeting minutes, interview compilation, interview review, and industry research. The product's templated output and to-do extraction capabilities can be embedded directly into workflows. Not suitable for the boundary: For internal meetings with a high level of confidentiality, using a SaaS platform to process recordings may involve data privacy compliance issues. In such scenarios, the enterprise's privatized deployment plan should be evaluated first.
-
Corporate/Institutional Purchasers: Customers from educational institutions, radio and television media, financial risk control, content review and other industries who need to process audio and video in batches. The enterprise version of the all-in-one solution supports 30 channels of concurrency and localized system adaptation, and is suitable for scenarios with strict requirements for data security and offline deployment. Unfit Boundary: In scenarios where there is strong demand for real-time live broadcast transcription (millisecond-level subtitles) and limited budget, the queuing processing mechanism of the SaaS version may not meet the latency requirements.
General unsuitability: Individual users with extremely poor audio quality (background noise > voice volume), high proportion of dense dialects or non-standard accents, strict requirements on real-time performance (such as real-time simultaneous interpretation), need to be completely offline, and cannot deploy the enterprise version.
Summary and Outlook
The core value of Multi-Dimensional Vision is to transform "watching videos" from passive consumption to active knowledge extraction. By covering 12 platform link compatibility + local file upload + full-link analysis template + content creation tools, it has formed a complete link in the "audio and video → structured knowledge" section. For individual users, its differentiation lies in the permanent validity of the analysis time, no mechanism to clear it at the end of the month, and pre-coverage of content creator scenarios (copywriting templates, subtitle burning).
Current Limitations and Uncertainties:
- Transcription ceiling: Transcription accuracy is highly dependent on audio quality, and its performance in complex acoustic environments has not yet been verified by independent benchmarks. For fields with dense professional terminology (law, medicine, engineering), it is recommended to use sample audio of the target scene for actual testing before purchasing.
- API Missing: The platform does not provide a public API and cannot embed third-party workflows or perform batch automation. The developer ecosystem has not yet been established.
- Low model transparency: The name and version of the underlying ASR/LLM model are not disclosed, and users cannot independently evaluate the iteration pace of technical capabilities.
- B-side pricing is not transparent: Enterprise version pricing relies entirely on business communication, and small and medium-sized teams cannot quickly evaluate procurement budgets.
- Privacy and Data Compliance: Audio/video files uploaded by the personal SaaS version are stored in the cloud, and the data retention period is 30 days (free) to the membership period (paid). Enterprise users involved in sensitive content should prioritize evaluating the enterprise's privatization deployment plan or confirm the confidentiality clause in the data processing agreement.
Procurement/Adoption Risk Assessment: The current applicable boundaries of Multidimensional Vision are clear - it is most suitable for "non-sensitive audio and video analysis" scenarios for individual creators and small and medium-sized teams. The verification cost is low (free 60 minutes), and you can try before you buy. For enterprise procurement, it is recommended to follow the four-step process of "free quota test quality → annual card small-scale pilot → enterprise version business negotiation → privatized deployment", focusing on verifying the transcription accuracy, concurrency performance and data confidentiality terms. Key observation points for the future of the product include: the opening timing and pricing of the API, the release of industry benchmark evaluation data, and the accumulation of enterprise version customer cases.
How to use multi-dimensional vision
- Web client: You can use it by visiting the official website and registering an account. Most functions do not require installation.
- API access: Provides RESTful API, developers can obtain the API Key and integrate it into their own applications.
Version Info
- Multidimensional Vision Online Version (current version) :Provide audio and video upload and link analysis in the form of an online platform, and output structured content. The official unified version number has not been disclosed, and there is no official precise date yet. It is recorded according to the current online form.
- Platform is online :The platform provides external audio and video intelligent analysis capabilities, supporting local file upload and mainstream platform link analysis. There is no official precise date yet, it is recorded according to the public collection time.
User Reviews