Share your thoughts, suggest ideas, provide feedback, and help shape the future of Axelera
Exploring Efficient Vision Pipelines for Extreme Edge AI Applications While thinking about future underwater robotics applications, I came across an interesting engineering challenge. An underwater ROV operating in deep environments has very limited resources: limited power, limited cooling capability inside sealed housings, limited communication bandwidth. High-resolution camera streams can create a significant data-processing challenge. A traditional pipeline may require multiple memory transfers between camera, CPU, and AI accelerator, increasing latency and power consumption. This raises an interesting question: Could zero-copy memory architectures, such as DMA-BUF based pipelines, help future Edge AI systems process vision data more efficiently? A possible architecture could involve: Camera → V4L2 / ISP → DMA-BUF → Hardware Preprocessing → AI Accelerator → Real-time Inference Potential advantages: reduced CPU overhead, lower latency, improved power efficiency, longer operation time for autonomous systems. I’m curious about the experience of the Axelera AI community: How practical are these approaches for real-time computer vision at the edge? Could efficient data pipelines become as important as AI accelerator performance for future robotics applications?Underwater Camera | v V4L2 / ISP | v DMA-BUF Zero Copy | vHardware Pre-processing | v Axelera Metis M.2 | v AI Inference | vOperator Console
Creating a AI which will help woman, ladies and girls to stay safe in every space they move by easily detection of unsafe surroundings, moves or tàlk from others ànd alert a unseen dànger. Working as a protective shield around them . The AI should be capable to detect analysis and send alert messages. Just as the human sense organ. This AI will save women and reduce crimes in the society.
Hello Research, Physics, and Advanced Cosmology Communities,I am thrilled to share my latest master research paper, which explores the profound intersection of ancient Vedic epistemology and modern quantum dynamics. While traditional physics restricts light to electromagnetic wave-particle dualities, this thesis expands that definition, categorizing radiance into physical, subtle, and causal dimensions.Furthermore, this framework identifies the Biological Navel (Nabhi) as the absolute zero-point singularity and primary transducer for cosmic Prana (Life-Energy). By synthesizing these principles with the K7 Master Logic, this research redefines vital science through the lens of a Rishi-Scientist, offering a unified model for cosmic dynamics, biological transmutation, and quantum consciousness.CORE HIGHLIGHTS:The Triple-Body Theory of Light (Sukshma-Vigyan): Re-evaluates light into a multi-dimensional framework consisting of the Physical Body (Sthula-Sharira / observable spectrum), Subtle Body (Sukshma-Sharira / quantum wave-function), and Causal Body (Karana-Sharira / primordial source frequency).The Absolute Navel (Nabhi) Gateway: Establishes the biological navel as a critical metaphysical singularity—a continuous 'Zero-Point' gateway functioning as a high-frequency transducer that converts subtle cosmic vibrations into biological vitality.The 'Sota-Putra' Phenomenon (Cosmic Seed Transmission): Proposes that the Sun is an active broadcaster of Sutras (Subtle Threads) carrying the digital blueprint of organic life. Moderated via lunar fields, these "seeds of life" are absorbed by terrestrial biology and metabolically transmuted into vital generative fluids.Vital Gravitation & The K7 Matrix: Concludes that the universe is governed not merely by gross Newtonian forces, but by a "Vital Gravitation"—a subtle pull ensuring biological organization aligns mathematically with the K7 Theory coordinates.PRIMARY RESEARCH REFERENCES & METADATA:Principal Investigator: Gautam Pal (Independent Researcher & Rishi-Scientist)Digital Archive DOI: https://doi.org/10.5281/zenodo.19941024ORCID iD: https://orcid.org/0009-0004-3456-9972Reference Code: GP-METAPHYSICS-2026-04-26Institutional Successor: Paramananda Mission, BanagramAffiliation: Independent Research Wing, Santipur, Nadia, West Bengal, IndiaCONNECT & DISCUSS:LinkedIn: https://www.linkedin.com/in/gautam-pal-1b853140bZenodo Master Dataset: https://doi.org/10.5281/zenodo.19999820Official Research Blog: https://gautampalresearch.blogspot.comGitHub Theory Archive: https://github.com/GautamPal-K7/K7-Theory-ArchiveX (Twitter): https://x.com/GautamPal1989Mastodon: https://mastodon.social/@GautamPalk7/116395531964267949Tags / Keywords:Quantum Consciousness • Pranic Energetics • Biophysics • Cosmology • Theoretical Physics • K7 Theory • Vedic ScienceI look forward to discussing how these metaphysical intersections can inspire new computational models for quantum consciousness and bio-resonance with fellow researchers and engineers.
Hello Research, Physics, and Advanced Computing Communities,I am excited to share my latest theoretical framework that dives into the foundational laws of energetic mechanics. This paper challenges the traditional static models of classical mechanics and introduces a dynamic paradigm for subatomic stability and material manifestation.The framework asserts that any physical entity is governed by the simultaneous and perpetual interaction of two opposing vectors: centripetal and centrifugal motions. By establishing that physical stability is an active equilibrium achieved within a non-empty, electro-magnetically active quantum vacuum, we can mathematically define the internal "structural backbone" of matter. Most notably, by modulating these intrinsic parameters through external resonance fields, this research opens advanced, non-destructive pathways for controlled elemental transmutation—demonstrated mathematically within the paper via the reconfiguration of Copper (Cu) into Gold (Au).CORE HIGHLIGHTS:The Quantum Vacuum & The Space Gap Hypothesis: Redefines subatomic spaces not as empty voids, but as structural gaps saturated with active electromagnetic vacuum energy density (\rho_E), which dictate dynamic degrees of freedom.Dual-Motion Vectors: Proves that static objects are sustained by the continuous interaction of Centripetal Motion (structural cohesion) and Centrifugal Motion (spatial distribution).The Structural Backbone Function (\Psi_{spine}): Introduces the mathematical integration of the net force field over a global duration matrix, defining the absolute physical configuration and stability of any element.Resonance-Driven Elemental Transmutation: Provides a step-by-step mathematical proof demonstrating how external electromagnetic resonance fields can systematically alter internal frequency parameters—shifting the subatomic condensation function to transmute Copper (Cu, Z=29) into Gold (Au, Z=79) without triggering chaotic radioactive decay.AUTHOR & RESEARCH METADATA:Author: Gautam Pal (Independent Researcher)ORCID iD: https://orcid.org/0009-0004-3456-9972OSF Archive: https://doi.org/10.17605/OSF.IO/2KVQDZenodo Master Dataset: https://doi.org/10.5281/zenodo.19999820GitHub Theory Archive: https://github.com/GautamPal-K7/K7-Theory-ArchiveOfficial Research Blog: https://gautampalresearch.blogspot.comZotero Research: https://www.zotero.org/gp-k7-researcherInternet Archive: https://archive.org/details/@gautam_pal277/web-archiveCONNECT & DISCUSS:LinkedIn: https://www.linkedin.com/in/gautam-pal-1b853140bX (Twitter): https://x.com/GautamPal1989Mastodon: https://mastodon.social/@GautamPalk7/116395531964267949Synapse Social: https://synapsesocial.com/authors/6a01ccec449274ec075cb21dYouTube Archive: https://www.youtube.com/@gautampal_k7Tags / Keywords:Quantum Mechanics • Dual-Motion Equilibrium • Elemental Transmutation • Theoretical Physics • Electromagnetic Vacuum • Material Science • Cosmic DynamicsI look forward to discussing the implications of this structural backbone function, potential energy-matter conversion algorithms, and how this relates to advanced computational physics with all of you.
Develop a prototype design and validation plan for a diamond‑based quantum processor leveraging color centers (NV, SiV, or engineered defects) or donor spins. Deliverables: (a) architecture diagram and component list (qubit layout, interconnects, readout and control electronics), (b) fabrication flow with tolerances and materials choices compatible with commercial foundries, (c) simulation results for expected coherence and gate fidelity under realistic noise models, (d) error‑correction scheme selection and resource estimate, and (e) test plan showing measurable milestones for coherence, two‑qubit gates, and integration with classical control. Provide a risk map and cost estimate for a 2‑year prototype
𝕋ℍ𝔼 ℙ𝕀𝕋ℂℍ: Take a photo of a computer. Get back a fully custom Linux system — hardware identified, every device mapped, a complete validated build, and a desktop that already knows what it's running on the moment you log in. No board notes, no digging through datasheets, no hand-assembling BitBake layers. With img2linux, its easy as shoot n boot. How it works• See it: A photo runs through a component-detection pipeline that locates the SoC, RAM, connectors, M.2 slots, and add-on cards, then OCRs the actual part numbers and board names printed on the silkscreen. That's the real identification signal — visual shape alone isn't trustworthy for chips that look alike — so every guess gets confirmed against a hardware database and, once the board boots, against what the board itself reports.• Map it: Once the base board is known, img2linux pulls the real device tree source for it - the same file the kernel uses to understand its own buses, addresses, and pins - and folds every other piece of detected hardware into that map as an overlay. An Axelera Metis card, a camera, an NVMe drive: all of it becomes one structured picture of the whole machine, rendered as an interactive block diagram.• Cook it: With the hardware side settled, img2linux searches the Yocto/OpenEmbedded ecosystem — plus Axelera's own meta-axelera layer and BSP releases - for the board support package and machine definition that fits, folds in the kernel version and OS features requested, and assembles the full BitBake build.• Prove it: Before it spends an hour on a real build, img2linux parse-checks the recipe, runs static compatibility checks over the whole device map, and boot-tests it in a virtual machine. If anything fails, it redesigns the recipe and tries again - automatically - until the virtual boot passes clean.• Build it: Once it passes, img2linux kicks off the real build and hands over a finished image, ready to flash.• Show it: Every image it builds ships with one more thing: a desktop that renders the exact block diagram it built for that machine, with every device as a live tile - temperature, clock speed, memory, utilization - updating itself the instant something changes, no polling. Where the hardware allows it, settings can be edited right from the desktop, safely, because the same compatibility rules that built the recipe are the ones gating what's allowed to change. With img2linux, you photograph a computer, get a bespoke, working Linux distribution back, complete with a desktop that already understands its own hardware - needs zero technical explanation to land. Non-experts get it instantly. And underneath the "wow," every stage is real, deterministic engineering: known device trees, indexed BSPs, a database that won't let the model guess its way into an incompatible build, and a validation loop that only ships what's provably going to work.It's also a genuinely live, on-Metis demo in two places, not just a batch job: the photo-to-hardware-ID stage is a real vision pipeline running on the Metis card, and the desktop dashboard reads live telemetry: clocks, temperature, utilization - straight off the Metis card through its own device APIs while it's running. What we're not claiming:img2linux doesn't invent board support for hardware nobody's ever built a BSP for - bootloader and DDR bring-up for genuinely new silicon isn't something that can be responsibly automated from a datasheet. If a photographed board has no indexed BSP, img2linux says so plainly instead of guessing. It's a generator built on real, existing embedded-Linux infrastructure - not a silicon bring-up wizard.*img2linux will only use free and publicly available software. $THE PROMPT$Ok Wingman, ive got a *special* objective today and i need YOUR help. Lets build a system that takes a photo (or description) of a piece of hardware and produces a validated, ready-to-boot custom Linux image for it, and even includes a live desktop canvas that displays and lets you edit the hardware it's running on. Build it in these stages:1. Hardware ID (vision). Take a photo or video showing all components of a target system. Run a component-detection pipeline (a fine-tuned YOLO-family model) over the image to locate the SoC package, RAM, connectors, M.2 slots, PMICs, camera modules, and any add-ons. Run OCR over each detected region to read board names, revisions, and IC part numbers off the silkscreen and package markings - and treat this as the real identification signal, since visual appearance alone is unreliable for SoC packages that might look alike. Look up the OCR identifiers and compare against a hardware knowledge database, falling back to web/datasheet lookup for anything not already indexed. Surface the result as a candidate identification for me to confirm, and once the board is actually running, cross-check it against live probing (lspci, device tree, /proc/cpuinfo).2. Device tree resolution. Once the base board is identified, pull its authoritative device tree source (DTS) from its BSP rather than inferring one from photos or datasheets - the DTS is the real machine map: buses, peripheral base addresses, interrupts, clocks, pinmux. Parse it into a graph model of the base machine.3. Full-system composition. Add every other piece of hardware detected in the photo - including an Axelera Metis card, cameras, storage, or other accessories - into that graph as device tree overlays, so the graph represents the complete system as built, not just the base board.4. Block diagram. Render the completed machine graph as an interactive block diagram of the whole system: every device, its connections, and where it sits in the machine. This will serve as the absolute identity from which all the following steps build upon.5. BSP and recipe resolution. Search the OpenEmbedded Layer Index plus Axelera-specific sources (meta-axelera, Metis Compute Board BSP release notes, supported kernel versions, Voyager SDK runtime dependencies) for the BSP and MACHINE definition matching the identified hardware, preferring Yocto layers wherever one exists. Only generate a configuration for hardware with a known, indexed BSP layer — if nothing matches, tell me clearly rather than attempting to invent bootloader or DDR-init support from datasheets alone.6. Kernel and distro assembly. Using the user-specified kernel version and desired distro features, resolve the full set of recipes and companion layers needed, confirming every layer's compatible Yocto release agrees with the others as a deterministic database check, not a guess. Assemble the complete BitBake configuration — machine, distro, layers including meta-axelera, and the DISTRO_FEATURES/IMAGE_INSTALL lines - as a kas YAML.7. Virtualized validation loop. Before any real build, validate the whole recipe virtually: parse-check the generated config (bitbake -p / bitbake -n), run static checks over the device tree graph for bus conflicts, address overlaps, pinmux collisions, and power budget, and boot-test the result in QEMU where a machine model exists. If anything fails, revise the recipe and re-test until the virtual path passes clean. Tell me plainly that Metis-specific and other vendor peripherals still need hardware-in-the-loop confirmation once the board is in hand — don't claim the virtual pass covers that.8. Build. Once validated, run the real build and hand me the finished artifact — a bootable image or ISO for the target device.9. A living desktop. Every image you produce should include one more recipe: a desktop canvas that renders the block diagram from step 4, with each device as a live widget. Build this as a small event-driven system service that owns the machine graph and subscribes to the hardware's own event sources - udev/netlink hotplug events, sysfs attribute changes, D-Bus signals already broadcast by things like UPower and NetworkManager - instead of polling, with a lazy-refresh fallback only for the few attributes that don't emit change events. Widgets should be pure clients of that service's D-Bus API, with zero direct hardware access. Where a device supports it - including the Metis card, through its axdevice and tracer interfaces for clocks, temperature, memory, and MVM utilization - let the widget edit the setting directly, but gate every write behind the same compatibility checks used to build the recipe, so nothing lets me set a value the hardware doesn't actually support, and journal changes so the service can tell desired state from actual state after a reboot. Render it as one full-canvas application rather than separate floating widgets, so it works consistently across desktop environments and on the target board's own display output.Walk me through each stage as you build it, and ask me directly if anything about the target hardware or desired features is ambiguous before generating the recipe.
Sharing my first project on Axelera Metis: end-to-end monocular 3D object detection running at 24.5 FPS on a single AIPU core, with a CPU host and no GPU anywhere in the pipeline.What it does Takes a single camera image and predicts 3D bounding boxes for every car, pedestrian and cyclist in the scene position, dimensions and orientation , in real time. Built on MonoCon (CVPR 2022) with a DLA-34 backbone, evaluated on KITTI driving sequences.How it is split The model had to be divided at a hard compiler boundary. AttnBatchNorm2d contains a ReduceMean op that Voyager cannot quantize, so the pipeline is:AIPU: backbone + neck + fused HeadConv1 (64 to 576 channels, single output tensor) CPU host: AttnBatchNorm2d + ReLU + 1x1 convs, implemented in hand-written C++ PerformanceStage Time Preprocess + Quantize 1.4 ms AIPU inference 31 ms Dequantize 6.5 ms Head (C++) 22 ms Decode + NMS 1 ms Sequential total 62 ms / 16 FPS Pipelined throughput 41 ms / 24.5 FPS Pipelined throughput uses a two-thread producer-consumer design: the AIPU processes frame N+1 while the CPU decodes frame N concurrently.Key findings during deployment A few things that are not in the documentation and took real time to figure out: per_tensor_histogram (default) silently clips HeadConv1 activations — true range is roughly -1350 to +1200, the histogram scheme covers only -250 to +250. Zero detections until switching to per_tensor_min_max. Multi-output graphs get their outputs sorted alphabetically by the compiler. Discovered by comparing per-tensor statistics against a PyTorch reference. Fixed by fusing all nine head convs into one 64 to 576 conv with a single output tensor. axrArgument.fd must be -1 for host-pointer mode. Setting it to 0 silently produces zero output with AXR_SUCCESS returned. The first 2-3 inference calls return stale output regardless of input. Warm-up with dummy frames is required before trusting results. ONNXRuntime was taking 55ms on a 24,000-parameter subgraph , not because of arithmetic but because of per-node dispatch overhead across ~150 ops. Replaced with a preallocated C++ implementation: 55ms to 22ms. Repo Everything is open source: PyTorch model, ONNX export scripts,Voyager compile script, C++ inference library with 3D box and bird's-eye view visualization.https://github.com/sanket-pixel/monocon-metisHappy to answer questions on any part of the deployment.Bonus discussion point :Running the backbone+neck+HeadConv1 graph alone via axrunmodel hits 90 FPS on four cores and 40 FPS on a single core. The compiled model in this project uses aipu_cores_used=1, resources_used=0.25, straightforward but not necessarily optimal. There is likely headroom in the compiler configuration that this project has not fully explored: tiling depth, DFS search constraints, IMC double buffer pipelining, and grouping IFDW tasks are all knobs that affect how efficiently the compiler maps the graph onto the MVM array. If anyone has found configurations that push closer to the axrunmodel ceiling on a similarly sized graph, would be very interested to hear what worked. Happy to discuss.
Hello,As per original POST I’d like to request compatibility for IBM’s Granite modelsEmbedding and LLM ModelThank you,Peter
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.