subscribe to arXiv mailings

Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation

Authors: Peng Jin, Hao Li, Zesen Cheng, Kehan Li, Runyi Yu, Chang Liu, Xiangyang Ji, Li Yuan, Jie Chen

Abstract: Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily focus on the direct synthesis of global motions while neglecting the importance of generating and controlling local actions. In this paper, we propose the local… ▽ More Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily focus on the direct synthesis of global motions while neglecting the importance of generating and controlling local actions. In this paper, we propose the local action-guided motion diffusion model, which facilitates global motion generation by utilizing local actions as fine-grained control signals. Specifically, we provide an automated method for reference local action sampling and leverage graph attention networks to assess the guiding weight of each local action in the overall motion synthesis. During the diffusion process for synthesizing global motion, we calculate the local-action gradient to provide conditional guidance. This local-to-global paradigm reduces the complexity associated with direct global motion generation and promotes motion diversity via sampling diverse actions as conditions. Extensive experiments on two human motion datasets, i.e., HumanML3D and KIT, demonstrate the effectiveness of our method. Furthermore, our method provides flexibility in seamlessly combining various local actions and continuous guiding weight adjustment, accommodating diverse user preferences, which may hold potential significance for the community. The project page is available at https://jpthu17.github.io/GuidedMotion-project/. △ Less

Submitted 15 July, 2024; originally announced July 2024.

Comments: Accepted by ECCV 2024

arXiv:2407.10424 [pdf, other]

CodeV: Empowering LLMs for Verilog Generation through Multi-Level Summarization

Authors: Yang Zhao, Di Huang, Chongxiao Li, Pengwei Jin, Ziyuan Nan, Tianyun Ma, Lei Qi, Yansong Pan, Zhenxing Zhang, Rui Zhang, Xishan Zhang, Zidong Du, Qi Guo, Xing Hu, Yunji Chen

Abstract: The increasing complexity and high costs associated with modern processor design have led to a surge in demand for processor design automation. Instruction-tuned large language models (LLMs) have demonstrated remarkable performance in automatically generating code for general-purpose programming languages like Python. However, these methods fail on hardware description languages (HDLs) like Verilo… ▽ More The increasing complexity and high costs associated with modern processor design have led to a surge in demand for processor design automation. Instruction-tuned large language models (LLMs) have demonstrated remarkable performance in automatically generating code for general-purpose programming languages like Python. However, these methods fail on hardware description languages (HDLs) like Verilog due to the scarcity of high-quality instruction tuning data, as even advanced LLMs like GPT-3.5 exhibit limited performance on Verilog generation. Regarding this issue, we observe that (1) Verilog code collected from the real world has higher quality than those generated by LLMs. (2) LLMs like GPT-3.5 excel in summarizing Verilog code rather than generating it. Based on these observations, this paper introduces CodeV, a series of open-source instruction-tuned Verilog generation LLMs. Instead of generating descriptions first and then getting the corresponding code from advanced LLMs, we prompt the LLM with Verilog code and let the LLM generate the corresponding natural language description by multi-level summarization. Experimental results show that CodeV relatively surpasses the previous open-source SOTA by 14.4% (BetterV in VerilogEval) and 11.3% (RTLCoder in RTLLM) respectively, and also relatively outperforms previous commercial SOTA GPT-4 by 22.1% in VerilogEval. △ Less

Submitted 15 July, 2024; v1 submitted 14 July, 2024; originally announced July 2024.

Comments: 16 pages, 8 figures, conference

arXiv:2407.08903 [pdf, other]

doi 10.1145/3622781.3674168

TensorTEE: Unifying Heterogeneous TEE Granularity for Efficient Secure Collaborative Tensor Computing

Authors: Husheng Han, Xinyao Zheng, Yuanbo Wen, Yifan Hao, Erhu Feng, Ling Liang, Jianan Mu, Xiaqing Li, Tianyun Ma, Pengwei Jin, Xinkai Song, Zidong Du, Qi Guo, Xing Hu

Abstract: Heterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is considered a promising solution because of its comparatively lower overhead. However, existing heterogeneous TEE designs are inefficient for collaborative computin… ▽ More Heterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is considered a promising solution because of its comparatively lower overhead. However, existing heterogeneous TEE designs are inefficient for collaborative computing due to fine and different memory granularities between CPU and NPU. 1) The cacheline granularity of CPU TEE intensifies memory pressure due to its extra memory access, and 2) the cacheline granularity MAC of NPU escalates the pressure on the limited memory storage. 3) Data transfer across heterogeneous enclaves relies on the transit of non-secure regions, resulting in cumbersome re-encryption and scheduling. To address these issues, we propose TensorTEE, a unified tensor-granularity heterogeneous TEE for efficient secure collaborative tensor computing. First, we virtually support tensor granularity in CPU TEE to eliminate the off-chip metadata access by detecting and maintaining tensor structures on-chip. Second, we propose tensor-granularity MAC management with predictive execution to avoid computational stalls while eliminating off-chip MAC storage and access. Moreover, based on the unified granularity, we enable direct data transfer without re-encryption and scheduling dilemmas. Our evaluation is built on enhanced Gem5 and a cycle-accurate NPU simulator. The results show that TensorTEE improves the performance of Large Language Model (LLM) training workloads by 4.0x compared to existing work and incurs only 2.1% overhead compared to non-secure training, offering a practical security assurance for LLM training. △ Less

Submitted 11 July, 2024; originally announced July 2024.

Comments: Accepted by ASPLOS 2024

arXiv:2407.04872 [pdf, ps, other]

Faster single-source shortest paths with negative real weights via proper hop distance

Authors: Yufan Huang, Peter Jin, Kent Quanrud

Abstract: The textbook algorithm for single-source shortest paths with real-valued edge weights runs in $O(m n)$ time on a graph with $m$ edges and $n$ vertices. A recent breakthrough algorithm by Fineman [Fin24] takes $\tilde O(m n^{8/9})$ randomized time. We present an $\tilde O(m n^{4/5})$ randomized time algorithm building on ideas from [Fin24]. The textbook algorithm for single-source shortest paths with real-valued edge weights runs in $O(m n)$ time on a graph with $m$ edges and $n$ vertices. A recent breakthrough algorithm by Fineman [Fin24] takes $\tilde O(m n^{8/9})$ randomized time. We present an $\tilde O(m n^{4/5})$ randomized time algorithm building on ideas from [Fin24]. △ Less

Submitted 5 July, 2024; originally announced July 2024.

arXiv:2407.04162 [pdf, other]

Measurement Embedded Schrödinger Bridge for Inverse Problems

Authors: Yuang Wang, Pengfei Jin, Siyeop Yoon, Matthew Tivnan, Quanzheng Li, Li Zhang, Dufan Wu

Abstract: Score-based diffusion models are frequently employed as structural priors in inverse problems. However, their iterative denoising process, initiated from Gaussian noise, often results in slow inference speeds. The Image-to-Image Schrödinger Bridge (I$^2$SB), which begins with the corrupted image, presents a promising alternative as a prior for addressing inverse problems. In this work, we introduc… ▽ More Score-based diffusion models are frequently employed as structural priors in inverse problems. However, their iterative denoising process, initiated from Gaussian noise, often results in slow inference speeds. The Image-to-Image Schrödinger Bridge (I$^2$SB), which begins with the corrupted image, presents a promising alternative as a prior for addressing inverse problems. In this work, we introduce the Measurement Embedded Schrödinger Bridge (MESB). MESB establishes Schrödinger Bridges between the distribution of corrupted images and the distribution of clean images given observed measurements. Based on optimal transport theory, we derive the forward and backward processes of MESB. Through validation on diverse inverse problems, our proposed approach exhibits superior performance compared to existing Schrödinger Bridge-based inverse problems solvers in both visual quality and quantitative metrics. △ Less

Submitted 22 May, 2024; originally announced July 2024.

Comments: 14 pages, 2 figures, Neurips preprint

arXiv:2406.18139 [pdf, other]

LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Authors: Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, Li Yuan

Abstract: Long-context Multimodal Large Language Models (MLLMs) demand substantial computational resources for inference as the growth of their multimodal Key-Value (KV) cache, in response to increasing input lengths, challenges memory and time efficiency. Unlike single-modality LLMs that manage only textual contexts, the KV cache of long-context MLLMs includes representations from multiple images with temp… ▽ More Long-context Multimodal Large Language Models (MLLMs) demand substantial computational resources for inference as the growth of their multimodal Key-Value (KV) cache, in response to increasing input lengths, challenges memory and time efficiency. Unlike single-modality LLMs that manage only textual contexts, the KV cache of long-context MLLMs includes representations from multiple images with temporal and spatial relationships and related textual contexts. The predominance of image tokens means traditional optimizations for LLMs' KV caches are unsuitable for multimodal long-context settings, and no prior works have addressed this challenge. In this work, we introduce LOOK-M, a pioneering, fine-tuning-free approach that efficiently reduces the multimodal KV cache size while maintaining performance comparable to a full cache. We observe that during prompt prefill, the model prioritizes more textual attention over image features, and based on the multimodal interaction observation, a new proposed text-prior method is explored to compress the KV cache. Furthermore, to mitigate the degradation of image contextual information, we propose several compensatory strategies using KV pairs merging. LOOK-M demonstrates that with a significant reduction in KV Cache memory usage, such as reducing it by 80% in some cases, it not only achieves up to 1.5x faster decoding but also maintains or even enhances performance across a variety of long context multimodal tasks. △ Less

Submitted 26 June, 2024; originally announced June 2024.

arXiv:2406.04481 [pdf, other]

Optimizing Autonomous Driving for Safety: A Human-Centric Approach with LLM-Enhanced RLHF

Authors: Yuan Sun, Navid Salami Pargoo, Peter J. Jin, Jorge Ortiz

Abstract: Reinforcement Learning from Human Feedback (RLHF) is popular in large language models (LLMs), whereas traditional Reinforcement Learning (RL) often falls short. Current autonomous driving methods typically utilize either human feedback in machine learning, including RL, or LLMs. Most feedback guides the car agent's learning process (e.g., controlling the car). RLHF is usually applied in the fine-t… ▽ More Reinforcement Learning from Human Feedback (RLHF) is popular in large language models (LLMs), whereas traditional Reinforcement Learning (RL) often falls short. Current autonomous driving methods typically utilize either human feedback in machine learning, including RL, or LLMs. Most feedback guides the car agent's learning process (e.g., controlling the car). RLHF is usually applied in the fine-tuning step, requiring direct human "preferences," which are not commonly used in optimizing autonomous driving models. In this research, we innovatively combine RLHF and LLMs to enhance autonomous driving safety. Training a model with human guidance from scratch is inefficient. Our framework starts with a pre-trained autonomous car agent model and implements multiple human-controlled agents, such as cars and pedestrians, to simulate real-life road environments. The autonomous car model is not directly controlled by humans. We integrate both physical and physiological feedback to fine-tune the model, optimizing this process using LLMs. This multi-agent interactive environment ensures safe, realistic interactions before real-world application. Finally, we will validate our model using data gathered from real-life testbeds located in New Jersey and New York City. △ Less

Submitted 6 June, 2024; originally announced June 2024.

arXiv:2406.02862 [pdf, other]

Rethinking Guidance Information to Utilize Unlabeled Samples:A Label Encoding Perspective

Authors: Yulong Zhang, Yuan Yao, Shuhao Chen, Pengrong Jin, Yu Zhang, Jian Jin, Jiangang Lu

Abstract: Empirical Risk Minimization (ERM) is fragile in scenarios with insufficient labeled samples. A vanilla extension of ERM to unlabeled samples is Entropy Minimization (EntMin), which employs the soft-labels of unlabeled samples to guide their learning. However, EntMin emphasizes prediction discriminability while neglecting prediction diversity. To alleviate this issue, in this paper, we rethink the… ▽ More Empirical Risk Minimization (ERM) is fragile in scenarios with insufficient labeled samples. A vanilla extension of ERM to unlabeled samples is Entropy Minimization (EntMin), which employs the soft-labels of unlabeled samples to guide their learning. However, EntMin emphasizes prediction discriminability while neglecting prediction diversity. To alleviate this issue, in this paper, we rethink the guidance information to utilize unlabeled samples. By analyzing the learning objective of ERM, we find that the guidance information for labeled samples in a specific category is the corresponding label encoding. Inspired by this finding, we propose a Label-Encoding Risk Minimization (LERM). It first estimates the label encodings through prediction means of unlabeled samples and then aligns them with their corresponding ground-truth label encodings. As a result, the LERM ensures both prediction discriminability and diversity, and it can be integrated into existing methods as a plugin. Theoretically, we analyze the relationships between LERM and ERM as well as EntMin. Empirically, we verify the superiority of the LERM under several label insufficient scenarios. The codes are available at https://github.com/zhangyl660/LERM. △ Less

Submitted 4 June, 2024; originally announced June 2024.

Comments: Accepted to ICML 2024

arXiv:2405.19465 [pdf, other]

RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter

Authors: Meng Cao, Haoran Tang, Jinfa Huang, Peng Jin, Can Zhang, Ruyang Liu, Long Chen, Xiaodan Liang, Li Yuan, Ge Li

Abstract: Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most state-of-the-art TVR methods learn image-to-video transfer learning based on large-scale pre-trained visionlanguage models (e.g., CLIP). However, fully fine-tuning these pre-trained models for TVR incurs prohibitively expensive computation costs. To this end, we propose to conduct efficient… ▽ More Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most state-of-the-art TVR methods learn image-to-video transfer learning based on large-scale pre-trained visionlanguage models (e.g., CLIP). However, fully fine-tuning these pre-trained models for TVR incurs prohibitively expensive computation costs. To this end, we propose to conduct efficient text-video Retrieval with a sparse-andcorrelated AdaPter (RAP), i.e., fine-tuning the pre-trained model with a few parameterized layers. To accommodate the text-video scenario, we equip our RAP with two indispensable characteristics: temporal sparsity and correlation. Specifically, we propose a low-rank modulation module to refine the per-image features from the frozen CLIP backbone, which accentuates salient frames within the video features while alleviating temporal redundancy. Besides, we introduce an asynchronous self-attention mechanism that first selects the top responsive visual patches and augments the correlation modeling between them with learnable temporal and patch offsets. Extensive experiments on four TVR datasets demonstrate that RAP achieves superior or comparable performance compared to the fully fine-tuned counterpart and other parameter-efficient fine-tuning methods. △ Less

Submitted 29 May, 2024; originally announced May 2024.

Comments: Accepted by ACL 2024 Findings

arXiv:2405.00002 [pdf, ps, other]

Reconfigurable nonreciprocal heat transport with natural bulk materials

Authors: Min Lei, Peng Jin, Liujun Xu, Jiping Huang

Abstract: Non-reciprocity is increasingly scrutinised in contemporary physics and engineering, especially in the realm of heat transport. This concept opens up novel avenues for directional heat transport and thermal regulation. Nonetheless, the development of non-reciprocal thermal metamaterials confronts three primary challenges: a constrained operational temperature range and structural scale, considerab… ▽ More Non-reciprocity is increasingly scrutinised in contemporary physics and engineering, especially in the realm of heat transport. This concept opens up novel avenues for directional heat transport and thermal regulation. Nonetheless, the development of non-reciprocal thermal metamaterials confronts three primary challenges: a constrained operational temperature range and structural scale, considerable energy dissipation, and the limitation of non-reciprocal effects due to external parameters that are not freely adjustable. To surmount these hurdles, we propose a reconfigurable approach to non-reciprocal heat transport. The design features a simple asymmetrical structure, with a flat panel made of natural, evenly-distributed material, positioned vertically on one side of a central horizontal board.We employ natural convection on the vertical plate to facilitate non-reciprocal heat transport. The reconfigurability of non-reciprocal heat transport is achieved by adjusting the number of vertical plates and their thermal conductivity. Theoretical computations of the rectification ratio are employed to quantify and forecast the non-reciprocal effect, corroborated by simulations and empirical studies. We also establish correlations between the rectification ratio and various parameters, including the vertical plates' positioning and height, ambient temperature, and the temperature differential between heat sources. Control over multiple parameters can efficaciously widen the scope of non-reciprocal control and streamline experimental procedures. Furthermore, our research exclusively utilises the natural convection of air for non-reciprocal heat transport, obviating the need for Supplementary energy sources and markedly enhancing energy efficiency. △ Less

Submitted 6 January, 2024; originally announced May 2024.

Comments: 34 pages, 15 figures

arXiv:2404.14125 [pdf, ps, other]

Weights for $π$-partial characters of $π$-separable groups

Authors: Xuewu Chang, Ping Jin

Abstract: The aim of this paper is to confirm an inequality predicted by Isaacs and Navarro in 1995, which asserts that for any $π'$-subgroup $Q$ of a $π$-separable group $G$, the number of $π'$-weights of $G$ with $Q$ as the first component always exceeds that of irreducible $π$-partial characters of $G$ with $Q$ as their vertex. We also give some sufficient condition to guarantee that these two numbers ar… ▽ More The aim of this paper is to confirm an inequality predicted by Isaacs and Navarro in 1995, which asserts that for any $π'$-subgroup $Q$ of a $π$-separable group $G$, the number of $π'$-weights of $G$ with $Q$ as the first component always exceeds that of irreducible $π$-partial characters of $G$ with $Q$ as their vertex. We also give some sufficient condition to guarantee that these two numbers are equal, and thereby strengthen their main theorem on the $π$-version of the Alperin weight conjecture. △ Less

Submitted 22 April, 2024; originally announced April 2024.

MSC Class: 20C15; 20C20

arXiv:2404.08916 [pdf, other]

Meply: A Large-scale Dataset and Baseline Evaluations for Metastatic Perirectal Lymph Node Detection and Segmentation

Authors: Weidong Guo, Hantao Zhang, Shouhong Wan, Bingbing Zou, Wanqin Wang, Chenyang Qiu, Jun Li, Peiquan Jin

Abstract: Accurate segmentation of metastatic lymph nodes in rectal cancer is crucial for the staging and treatment of rectal cancer. However, existing segmentation approaches face challenges due to the absence of pixel-level annotated datasets tailored for lymph nodes around the rectum. Additionally, metastatic lymph nodes are characterized by their relatively small size, irregular shapes, and lower contra… ▽ More Accurate segmentation of metastatic lymph nodes in rectal cancer is crucial for the staging and treatment of rectal cancer. However, existing segmentation approaches face challenges due to the absence of pixel-level annotated datasets tailored for lymph nodes around the rectum. Additionally, metastatic lymph nodes are characterized by their relatively small size, irregular shapes, and lower contrast compared to the background, further complicating the segmentation task. To address these challenges, we present the first large-scale perirectal metastatic lymph node CT image dataset called Meply, which encompasses pixel-level annotations of 269 patients diagnosed with rectal cancer. Furthermore, we introduce a novel lymph-node segmentation model named CoSAM. The CoSAM utilizes sequence-based detection to guide the segmentation of metastatic lymph nodes in rectal cancer, contributing to improved localization performance for the segmentation model. It comprises three key components: sequence-based detection module, segmentation module, and collaborative convergence unit. To evaluate the effectiveness of CoSAM, we systematically compare its performance with several popular segmentation methods using the Meply dataset. Our code and dataset will be publicly available at: https://github.com/kanydao/CoSAM. △ Less

Submitted 13 April, 2024; originally announced April 2024.

Comments: 13 pages

arXiv:2404.08446 [pdf]

Growth of two-inch free-standing heteroepitaxial diamond on Ir/YSZ/Si (001) substrates via laser-patterned templates

Authors: Pengfei Qu, Peng Jin, Guangdi Zhou, Zhen Wang, Zhanguo Wang

Abstract: In this paper, 2-inch free-standing diamonds were prepared by using heteroepitaxy on composite Ir/YSZ/Si (001) substrates. To release stress, patterned templates were fabricated using laser etching after the initial growth of 50-nm-diamond. Then, the subsequent growth was completed on a patterned template. The full width at half maximum of the diamond (400) and (311) X-ray rocking curves were 313.… ▽ More In this paper, 2-inch free-standing diamonds were prepared by using heteroepitaxy on composite Ir/YSZ/Si (001) substrates. To release stress, patterned templates were fabricated using laser etching after the initial growth of 50-nm-diamond. Then, the subsequent growth was completed on a patterned template. The full width at half maximum of the diamond (400) and (311) X-ray rocking curves were 313.5 and 359.3 arcsecs, respectively. Strong band-edge emission in the cathodoluminescence spectrum of the resulting diamond revealed excellent crystalline quality. Furthermore, the 2D mapping of Raman spectra was conducted on a $2 mm \times 2 mm$ area located at the center of the 2-inch sample with a thickness of $400 μm$. The result showed an average peak width of $2.85 \pm 0.36 cm^{-1}$ and residual stress of $-0.03 \pm 0.37 GPa$. The dislocation density, determined by counting etching pits generated from $ H_2/O_2$ plasma etching, was estimated to be around $2.2 \times 10^7 cm^{-2}$. These results evidence that the laser-patterned method can effectively release stress during the growth of large-size diamonds, offering a simpler and more cost-effective alternative to the traditional photolithography-patterned scheme. △ Less

Submitted 12 April, 2024; originally announced April 2024.

Comments: 13 pages, 5 figures

arXiv:2403.06069 [pdf, other]

Implicit Image-to-Image Schrodinger Bridge for CT Super-Resolution and Denoising

Authors: Yuang Wang, Siyeop Yoon, Pengfei Jin, Matthew Tivnan, Zhennong Chen, Rui Hu, Li Zhang, Zhiqiang Chen, Quanzheng Li, Dufan Wu

Abstract: Conditional diffusion models have gained recognition for their effectiveness in image restoration tasks, yet their iterative denoising process, starting from Gaussian noise, often leads to slow inference speeds. As a promising alternative, the Image-to-Image Schrödinger Bridge (I2SB) initializes the generative process from corrupted images and integrates training techniques from conditional diffus… ▽ More Conditional diffusion models have gained recognition for their effectiveness in image restoration tasks, yet their iterative denoising process, starting from Gaussian noise, often leads to slow inference speeds. As a promising alternative, the Image-to-Image Schrödinger Bridge (I2SB) initializes the generative process from corrupted images and integrates training techniques from conditional diffusion models. In this study, we extended the I2SB method by introducing the Implicit Image-to-Image Schrodinger Bridge (I3SB), transitioning its generative process to a non-Markovian process by incorporating corrupted images in each generative step. This enhancement empowers I3SB to generate images with better texture restoration using a small number of generative steps. The proposed method was validated on CT super-resolution and denoising tasks and outperformed existing methods, including the conditional denoising diffusion probabilistic model (cDDPM) and I2SB, in both visual quality and quantitative metrics. These findings underscore the potential of I3SB in improving medical image restoration by providing fast and accurate generative modeling. △ Less

Submitted 9 March, 2024; originally announced March 2024.

arXiv:2403.05809 [pdf, other]

Shallow ReLU neural networks and finite elements

Authors: Pengzhan Jin

Abstract: We point out that (continuous or discontinuous) piecewise linear functions on a convex polytope mesh can be represented by two-hidden-layer ReLU neural networks in a weak sense. In addition, the numbers of neurons of the two hidden layers required to weakly represent are accurately given based on the numbers of polytopes and hyperplanes involved in this mesh. The results naturally hold for constan… ▽ More We point out that (continuous or discontinuous) piecewise linear functions on a convex polytope mesh can be represented by two-hidden-layer ReLU neural networks in a weak sense. In addition, the numbers of neurons of the two hidden layers required to weakly represent are accurately given based on the numbers of polytopes and hyperplanes involved in this mesh. The results naturally hold for constant and linear finite element functions. Such weak representation establishes a bridge between shallow ReLU neural networks and finite element functions, and leads to a perspective for analyzing approximation capability of ReLU neural networks in $L^p$ norm via finite element functions. Moreover, we discuss the strict representation for tensor finite element functions via the recent tensor neural networks. △ Less

Submitted 9 March, 2024; originally announced March 2024.

arXiv:2403.02874 [pdf, other]

The bright black hole X-ray binary 4U 1543-47 during 2021 outburst. A clear state transition from super-Eddington to sub-Eddington accretion revealed by Insight-HXMT

Authors: Pei Jin, Guobao Zhang, Yuexin Zhang, Mariano Méndez, Jinlu Qu, David M. Russell, Jiancheng Wang, Shuangnan Zhang, Yi-Jung Yang, Shumei Jia, Zixu Yang, Hexin Liu

Abstract: We present a detailed analysis of the observations with the Hard X-ray Modulation Telescope of the black hole X-ray transient 4U~1543-47 during its outburst in 2021. We find a clear state transition during the outburst decay of the source. Using previous measurements of the black-hole mass and distance to the source, the source luminosity during this transition is close to the Eddington limit. The… ▽ More We present a detailed analysis of the observations with the Hard X-ray Modulation Telescope of the black hole X-ray transient 4U~1543-47 during its outburst in 2021. We find a clear state transition during the outburst decay of the source. Using previous measurements of the black-hole mass and distance to the source, the source luminosity during this transition is close to the Eddington limit. The light curves before and after the transition can be fitted by two exponential functions with short ($\sim 16$ days) and long ($\sim 130$ days) decay time scales, respectively. We detect strong reflection features in all observations that can be described with either the RelxillNS or Reflionx_bb reflection models, both of which have a black-body incident spectrum. In the super-Eddington state, we observe a Comptonized component characterized by a low electron temperature of approximately 2.0 keV. We suggest that this component appears exclusively within the inner radiation-pressure dominated region of the supercritical disk as a part of the intrinsic spectrum of the accretion disk itself. This feature vanishes as the source transitions into the sub-Eddington state. The emissivity index of the accretion disk in the reflection component is significantly different before and after the transition, $\sim3.0$-$5.0$ and $\sim7.0$-$9.0$ in the super- and sub-Eddington states, respectively. Based on the reflection geometry of returning disk radiation, the geometrically thicker the accretion disk, the smaller the emissivity index. Therefore, we propose that the transition is primarily driven by the change of the accretion flow from a supercritical to a thin disk configuration. △ Less

Submitted 5 March, 2024; originally announced March 2024.

arXiv:2402.15097 [pdf, other]

Learning solution operators of PDEs defined on varying domains via MIONet

Authors: Shanshan Xiao, Pengzhan Jin, Yifa Tang

Abstract: In this work, we propose a method to learn the solution operators of PDEs defined on varying domains via MIONet, and theoretically justify this method. We first extend the approximation theory of MIONet to further deal with metric spaces, establishing that MIONet can approximate mappings with multiple inputs in metric spaces. Subsequently, we construct a set consisting of some appropriate regions… ▽ More In this work, we propose a method to learn the solution operators of PDEs defined on varying domains via MIONet, and theoretically justify this method. We first extend the approximation theory of MIONet to further deal with metric spaces, establishing that MIONet can approximate mappings with multiple inputs in metric spaces. Subsequently, we construct a set consisting of some appropriate regions and provide a metric on this set thus make it a metric space, which satisfies the approximation condition of MIONet. Building upon the theoretical foundation, we are able to learn the solution mapping of a PDE with all the parameters varying, including the parameters of the differential operator, the right-hand side term, the boundary condition, as well as the domain. Without loss of generality, we for example perform the experiments for 2-d Poisson equations, where the domains and the right-hand side terms are varying. The results provide insights into the performance of this method across convex polygons, polar regions with smooth boundary, and predictions for different levels of discretization on one task. We also show the additional result of the fully-parameterized case in the appendix for interested readers. Reasonably, we point out that this is a meshless method, hence can be flexibly used as a general solver for a type of PDE. △ Less

Submitted 16 March, 2024; v1 submitted 23 February, 2024; originally announced February 2024.

arXiv:2402.14891 [pdf, other]

LLMBind: A Unified Modality-Task Integration Framework

Authors: Bin Zhu, Munan Ning, Peng Jin, Bin Lin, Jinfa Huang, Qi Song, Junwu Zhang, Zhenyu Tang, Mingjun Pan, Xing Zhou, Li Yuan

Abstract: In the multi-modal domain, the dependence of various models on specific input formats leads to user confusion and hinders progress. To address this challenge, we introduce \textbf{LLMBind}, a novel framework designed to unify a diverse array of multi-modal tasks. By harnessing a Mixture-of-Experts (MoE) Large Language Model (LLM), LLMBind processes multi-modal inputs and generates task-specific to… ▽ More In the multi-modal domain, the dependence of various models on specific input formats leads to user confusion and hinders progress. To address this challenge, we introduce \textbf{LLMBind}, a novel framework designed to unify a diverse array of multi-modal tasks. By harnessing a Mixture-of-Experts (MoE) Large Language Model (LLM), LLMBind processes multi-modal inputs and generates task-specific tokens, enabling the invocation of corresponding models to accomplish tasks. This unique approach empowers LLMBind to interpret inputs and generate outputs across various modalities, including image, text, video, and audio. Furthermore, we have constructed an interaction dataset comprising 400k instructions, which unlocks the ability of LLMBind for interactive visual generation and editing tasks. Extensive experimentation demonstrates that LLMBind achieves very superior performance across diverse tasks and outperforms existing models in user evaluations conducted in real-world scenarios. Moreover, the adaptability of LLMBind allows for seamless integration with the latest models and extension to new modality tasks, highlighting its potential to serve as a unified AI agent for modeling universal modalities. △ Less

Submitted 18 April, 2024; v1 submitted 22 February, 2024; originally announced February 2024.

arXiv:2402.07156 [pdf, other]

A hybrid iterative method based on MIONet for PDEs: Theory and numerical examples

Authors: Jun Hu, Pengzhan Jin

Abstract: We propose a hybrid iterative method based on MIONet for PDEs, which combines the traditional numerical iterative solver and the recent powerful machine learning method of neural operator, and further systematically analyze its theoretical properties, including the convergence condition, the spectral behavior, as well as the convergence rate, in terms of the errors of the discretization and the mo… ▽ More We propose a hybrid iterative method based on MIONet for PDEs, which combines the traditional numerical iterative solver and the recent powerful machine learning method of neural operator, and further systematically analyze its theoretical properties, including the convergence condition, the spectral behavior, as well as the convergence rate, in terms of the errors of the discretization and the model inference. We show the theoretical results for the frequently-used smoothers, i.e. Richardson (damped Jacobi) and Gauss-Seidel. We give an upper bound of the convergence rate of the hybrid method w.r.t. the model correction period, which indicates a minimum point to make the hybrid iteration converge fastest. Several numerical examples including the hybrid Richardson (Gauss-Seidel) iteration for the 1-d (2-d) Poisson equation are presented to verify our theoretical results, and also reflect an excellent acceleration effect. As a meshless acceleration method, it is provided with enormous potentials for practice applications. △ Less

Submitted 11 February, 2024; originally announced February 2024.

arXiv:2402.05935 [pdf, other]

SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Authors: Dongyang Liu, Renrui Zhang, Longtian Qiu, Siyuan Huang, Weifeng Lin, Shitian Zhao, Shijie Geng, Ziyi Lin, Peng Jin, Kaipeng Zhang, Wenqi Shao, Chao Xu, Conghui He, Junjun He, Hao Shao, Pan Lu, Hongsheng Li, Yu Qiao, Peng Gao

Abstract: We propose SPHINX-X, an extensive Multimodality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we… ▽ More We propose SPHINX-X, an extensive Multimodality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multimodal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral8x7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory △ Less

Submitted 26 June, 2024; v1 submitted 8 February, 2024; originally announced February 2024.

Comments: Accepted by ICML 2024. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory

arXiv:2401.15947 [pdf, other]

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Authors: Bin Lin, Zhenyu Tang, Yang Ye, Jiaxi Cui, Bin Zhu, Peng Jin, Jinfa Huang, Junwu Zhang, Yatian Pang, Munan Ning, Li Yuan

Abstract: Recent advances demonstrate that scaling Large Vision-Language Models (LVLMs) effectively improves downstream task performances. However, existing scaling methods enable all model parameters to be active for each token in the calculation, which brings massive training and inferring costs. In this work, we propose a simple yet effective training strategy MoE-Tuning for LVLMs. This strategy innovati… ▽ More Recent advances demonstrate that scaling Large Vision-Language Models (LVLMs) effectively improves downstream task performances. However, existing scaling methods enable all model parameters to be active for each token in the calculation, which brings massive training and inferring costs. In this work, we propose a simple yet effective training strategy MoE-Tuning for LVLMs. This strategy innovatively addresses the common issue of performance degradation in multi-modal sparsity learning, consequently constructing a sparse model with an outrageous number of parameters but a constant computational cost. Furthermore, we present the MoE-LLaVA, a MoE-based sparse LVLM architecture, which uniquely activates only the top-k experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments show the significant performance of MoE-LLaVA in a variety of visual understanding and object hallucination benchmarks. Remarkably, with only approximately 3B sparsely activated parameters, MoE-LLaVA demonstrates performance comparable to the LLaVA-1.5-7B on various visual understanding datasets and even surpasses the LLaVA-1.5-13B in object hallucination benchmark. Through MoE-LLaVA, we aim to establish a baseline for sparse LVLMs and provide valuable insights for future research in developing more efficient and effective multi-modal learning systems. Code is released at https://github.com/PKU-YuanGroup/MoE-LLaVA. △ Less

Submitted 6 July, 2024; v1 submitted 29 January, 2024; originally announced January 2024.

Comments: K = P + N represents the length of the output sequence in the formula (8)

arXiv:2401.09732 [pdf, other]

Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation

Authors: Zesen Cheng, Kehan Li, Hao Li, Peng Jin, Chang Liu, Xiawu Zheng, Rongrong Ji, Jie Chen

Abstract: Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage image-text pretraining model for recognizing object instances by separately aligning each frame and class texts, ignoring the correlation between frames. As a result, the separation breaks… ▽ More Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage image-text pretraining model for recognizing object instances by separately aligning each frame and class texts, ignoring the correlation between frames. As a result, the separation breaks the instance movement context of videos, causing inferior alignment between video and text. To tackle this issue, we propose to link frame-level instance representations as a Brownian Bridge to model instance dynamics and align bridge-level instance representation to class texts for more precisely open-vocabulary VIS (BriVIS). Specifically, we build our system upon a frozen video segmentor to generate frame-level instance queries, and design Temporal Instance Resampler (TIR) to generate queries with temporal context from frame queries. To mold instance queries to follow Brownian bridge and accomplish alignment with class texts, we design Bridge-Text Alignment (BTA) to learn discriminative bridge-level representations of instances via contrastive objectives. Setting MinVIS as the basic video segmentor, BriVIS surpasses the Open-vocabulary SOTA (OV2Seg) by a clear margin. For example, on the challenging large-vocabulary VIS dataset (BURST), BriVIS achieves 7.43 mAP and exhibits 49.49% improvement compared to OV2Seg (4.97 mAP). △ Less

Submitted 18 January, 2024; originally announced January 2024.

arXiv:2401.04148 [pdf, other]

Online Test-Time Adaptation of Spatial-Temporal Traffic Flow Forecasting

Authors: Pengxin Guo, Pengrong Jin, Ziyue Li, Lei Bai, Yu Zhang

Abstract: Accurate spatial-temporal traffic flow forecasting is crucial in aiding traffic managers in implementing control measures and assisting drivers in selecting optimal travel routes. Traditional deep-learning based methods for traffic flow forecasting typically rely on historical data to train their models, which are then used to make predictions on future data. However, the performance of the traine… ▽ More Accurate spatial-temporal traffic flow forecasting is crucial in aiding traffic managers in implementing control measures and assisting drivers in selecting optimal travel routes. Traditional deep-learning based methods for traffic flow forecasting typically rely on historical data to train their models, which are then used to make predictions on future data. However, the performance of the trained model usually degrades due to the temporal drift between the historical and future data. To make the model trained on historical data better adapt to future data in a fully online manner, this paper conducts the first study of the online test-time adaptation techniques for spatial-temporal traffic flow forecasting problems. To this end, we propose an Adaptive Double Correction by Series Decomposition (ADCSD) method, which first decomposes the output of the trained model into seasonal and trend-cyclical parts and then corrects them by two separate modules during the testing phase using the latest observed data entry by entry. In the proposed ADCSD method, instead of fine-tuning the whole trained model during the testing phase, a lite network is attached after the trained model, and only the lite network is fine-tuned in the testing process each time a data entry is observed. Moreover, to satisfy that different time series variables may have different levels of temporal drift, two adaptive vectors are adopted to provide different weights for different time series variables. Extensive experiments on four real-world traffic flow forecasting datasets demonstrate the effectiveness of the proposed ADCSD method. The code is available at https://github.com/Pengxin-Guo/ADCSD. △ Less

Submitted 8 January, 2024; originally announced January 2024.

arXiv:2401.03364 [pdf, other]

A dynamic thermal sensing mechanism with reconfigurable expanded-plane structures

Authors: Haohan Tan, Haoyang Cai, Peng Jin, Jiping Huang

Abstract: The precise measurement of temperature is crucial in various fields such as biology, medicine, industrial automation, energy management, and daily life applications. While in most scenarios, sensors with a fixed thermal conductivity inevitably mismatch the analogous parameter of the medium being measured, thus causing the distortion and inaccurate detection of original temperature fields. Despite… ▽ More The precise measurement of temperature is crucial in various fields such as biology, medicine, industrial automation, energy management, and daily life applications. While in most scenarios, sensors with a fixed thermal conductivity inevitably mismatch the analogous parameter of the medium being measured, thus causing the distortion and inaccurate detection of original temperature fields. Despite recent efforts on addressing the parameter-mismatch issue, all current solutions are constrained to a fixed working medium whereas a more universal sensor should function in a variety of scenes. Here, we report a dynamic thermal sensor capable of highly accurate measurements in diverse working environments. Remarkably, thanks to the highly tunable thermal conductivity of the expanded-plane structure, this sensor works effect on background mediums with a wide range of conductivity. Such a development greatly enhances the robustness and adaptability of thermal sensors, setting a solid foundation for applications in multi-physical sensing scenarios. △ Less

Submitted 6 January, 2024; originally announced January 2024.

arXiv:2312.17634 [pdf, other]

doi 10.3390/s24031021

Developing Flying Explorer for Autonomous Digital Modelling in Wild Unknowns

Authors: Naizhong Zhang. Yaoqiang Pan, Yangwen Jin, Peiqi Jin, Kewei Hu, Xiao Huang, Hanwen Kang

Abstract: This work presents an innovative solution for robotic odometry, path planning and exploration in wild unknown environments, focusing on digital modelling. The approach uses a minimum cost formulation with pseudo-randomly generated objectives, integrating multi-path planning and evaluation, with emphasis on full coverage of unknown maps based on feasible boundaries of interest. The evaluation carri… ▽ More This work presents an innovative solution for robotic odometry, path planning and exploration in wild unknown environments, focusing on digital modelling. The approach uses a minimum cost formulation with pseudo-randomly generated objectives, integrating multi-path planning and evaluation, with emphasis on full coverage of unknown maps based on feasible boundaries of interest. The evaluation carried out on a robotic platform with a lightweight 3D LiDAR sensor model, assesses the consistency and efficiency in exploring completely unknown subterranean-like areas. The algorithm allows for dynamic changes to the desired target and behaviour. At the same time, the paper details the design of AREX, highlighting its robust localisation, mapping and efficient exploration target selection capabilities, with a focus on continuity in exploration direction for increased efficiency and reduced odometry errors. The real-time, high-precision environmental perception module is identified as critical for accurate obstacle avoidance and exploration boundary identification. △ Less

Submitted 29 December, 2023; originally announced December 2023.

arXiv:2312.13271 [pdf, other]

Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting

Authors: Junwu Zhang, Zhenyu Tang, Yatian Pang, Xinhua Cheng, Peng Jin, Yida Wei, Munan Ning, Li Yuan

Abstract: Recent one image to 3D generation methods commonly adopt Score Distillation Sampling (SDS). Despite the impressive results, there are multiple deficiencies including multi-view inconsistency, over-saturated and over-smoothed textures, as well as the slow generation speed. To address these deficiencies, we present Repaint123 to alleviate multi-view bias as well as texture degradation and speed up t… ▽ More Recent one image to 3D generation methods commonly adopt Score Distillation Sampling (SDS). Despite the impressive results, there are multiple deficiencies including multi-view inconsistency, over-saturated and over-smoothed textures, as well as the slow generation speed. To address these deficiencies, we present Repaint123 to alleviate multi-view bias as well as texture degradation and speed up the generation process. The core idea is to combine the powerful image generation capability of the 2D diffusion model and the texture alignment ability of the repainting strategy for generating high-quality multi-view images with consistency. We further propose visibility-aware adaptive repainting strength for overlap regions to enhance the generated image quality in the repainting process. The generated high-quality and multi-view consistent images enable the use of simple Mean Square Error (MSE) loss for fast 3D content generation. We conduct extensive experiments and show that our method has a superior ability to generate high-quality 3D content with multi-view consistency and fine textures in 2 minutes from scratch. Our project page is available at https://pku-yuangroup.github.io/repaint123/. △ Less

Submitted 27 December, 2023; v1 submitted 20 December, 2023; originally announced December 2023.

Comments: Project page: https://pku-yuangroup.github.io/repaint123/

arXiv:2312.02428 [pdf, other]

FreestyleRet: Retrieving Images from Style-Diversified Queries

Authors: Hao Li, Curise Jia, Peng Jin, Zesen Cheng, Kehan Li, Jialu Sui, Chang Liu, Li Yuan

Abstract: Image Retrieval aims to retrieve corresponding images based on a given query. In application scenarios, users intend to express their retrieval intent through various query styles. However, current retrieval tasks predominantly focus on text-query retrieval exploration, leading to limited retrieval query options and potential ambiguity or bias in user intention. In this paper, we propose the Style… ▽ More Image Retrieval aims to retrieve corresponding images based on a given query. In application scenarios, users intend to express their retrieval intent through various query styles. However, current retrieval tasks predominantly focus on text-query retrieval exploration, leading to limited retrieval query options and potential ambiguity or bias in user intention. In this paper, we propose the Style-Diversified Query-Based Image Retrieval task, which enables retrieval based on various query styles. To facilitate the novel setting, we propose the first Diverse-Style Retrieval dataset, encompassing diverse query styles including text, sketch, low-resolution, and art. We also propose a light-weighted style-diversified retrieval framework. For various query style inputs, we apply the Gram Matrix to extract the query's textural features and cluster them into a style space with style-specific bases. Then we employ the style-init prompt tuning module to enable the visual encoder to comprehend the texture and style information of the query. Experiments demonstrate that our model, employing the style-init prompt tuning strategy, outperforms existing retrieval models on the style-diversified retrieval task. Moreover, style-diversified queries~(sketch+text, art+text, etc) can be simultaneously retrieved in our model. The auxiliary information from other queries enhances the retrieval performance within the respective query. △ Less

Submitted 8 December, 2023; v1 submitted 4 December, 2023; originally announced December 2023.

Comments: 16 pages, 7 figures

arXiv:2311.10122 [pdf, other]

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Authors: Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, Li Yuan

Abstract: The Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models. However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging f… ▽ More The Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models. However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging for a Large Language Model (LLM) to learn multi-modal interactions from several poor projection layers. In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM. As a result, we establish a simple but robust LVLM baseline, Video-LLaVA, which learns from a mixed dataset of images and videos, mutually enhancing each other. Video-LLaVA achieves superior performances on a broad range of 9 image benchmarks across 5 image question-answering datasets and 4 image benchmark toolkits. Additionally, our Video-LLaVA also outperforms Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively. Notably, extensive experiments demonstrate that Video-LLaVA mutually benefits images and videos within a unified visual representation, outperforming models designed specifically for images or videos. We aim for this work to provide modest insights into the multi-modal inputs for the LLM. △ Less

Submitted 21 November, 2023; v1 submitted 16 November, 2023; originally announced November 2023.

arXiv:2311.08046 [pdf, other]

Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Authors: Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, Li Yuan

Abstract: Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter challenges in effectively handling both image and video understanding, particularly with limited visual tokens. In this work, we introduce Chat-UniVi, a Unified Vision-language mo… ▽ More Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter challenges in effectively handling both image and video understanding, particularly with limited visual tokens. In this work, we introduce Chat-UniVi, a Unified Vision-language model capable of comprehending and engaging in conversations involving images and videos through a unified visual representation. Specifically, we employ a set of dynamic visual tokens to uniformly represent images and videos. This representation framework empowers the model to efficiently utilize a limited number of visual tokens to simultaneously capture the spatial details necessary for images and the comprehensive temporal relationship required for videos. Moreover, we leverage a multi-scale representation, enabling the model to perceive both high-level semantic concepts and low-level visual details. Notably, Chat-UniVi is trained on a mixed dataset containing both images and videos, allowing direct application to tasks involving both mediums without requiring any modifications. Extensive experimental results demonstrate that Chat-UniVi consistently outperforms even existing methods exclusively designed for either images or videos. Code is available at https://github.com/PKU-YuanGroup/Chat-UniVi. △ Less

Submitted 5 April, 2024; v1 submitted 14 November, 2023; originally announced November 2023.

Comments: Accepted by CVPR 2024 (Highlight)

arXiv:2311.01015 [pdf, other]

Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic Graphs

Authors: Peng Jin, Yang Wu, Yanbo Fan, Zhongqian Sun, Yang Wei, Li Yuan

Abstract: Most text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly di… ▽ More Most text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-training weights are available at https://github.com/jpthu17/GraphMotion. △ Less

Submitted 2 November, 2023; originally announced November 2023.

Comments: Accepted by NeurIPS 2023

arXiv:2309.13282 [pdf, other]

Convective Heat Transfer in Porous Materials

Authors: Peng Jin, Gaole Dai, Fubao Yang

Abstract: Thermal convection stands out as an exceptionally efficient thermal transport mechanism, distinctly separate from conduction and radiation. Yet, the inherently elusive nature of fluid motion poses challenges in accurately controlling convective heat flow. While recent innovations have harnessed thermal convection to achieve effective thermal conductivity, fusing thermal convection in liquids and t… ▽ More Thermal convection stands out as an exceptionally efficient thermal transport mechanism, distinctly separate from conduction and radiation. Yet, the inherently elusive nature of fluid motion poses challenges in accurately controlling convective heat flow. While recent innovations have harnessed thermal convection to achieve effective thermal conductivity, fusing thermal convection in liquids and thermal conduction in solids together to form hybrid thermal metamaterials is still challenging. In this review, we introduce the latest progress in convective heat transfer. Leveraging the right porous materials as a medium allows for a harmonious balance and synergy between convection and conduction, establishing stable heat and fluid flows. This paves the way for the innovative advancements in transformation thermotics. These findings demonstrate the remarkable tunability of convective heat transport in complex multicomponent thermal metamaterials. △ Less

Submitted 23 September, 2023; originally announced September 2023.

arXiv:2309.04711 [pdf, other]

doi 10.1103/RevModPhys.96.015002

Controlling mass and energy diffusion with metamaterials

Authors: Fubao Yang, Zeren Zhang, Liujun Xu, Zhoufei Liu, Peng Jin, Pengfei Zhuang, Min Lei, Jinrong Liu, Jian-Hua Jiang, Xiaoping Ouyang, Fabio Marchesoni, Jiping Huang

Abstract: Diffusion driven by temperature or concentration gradients is a fundamental mechanism of energy and mass transport, which inherently differs from wave propagation in both physical foundations and application prospects. Compared with conventional schemes, metamaterials provide an unprecedented potential for governing diffusion processes, based on emerging theories like the transformation and the sc… ▽ More Diffusion driven by temperature or concentration gradients is a fundamental mechanism of energy and mass transport, which inherently differs from wave propagation in both physical foundations and application prospects. Compared with conventional schemes, metamaterials provide an unprecedented potential for governing diffusion processes, based on emerging theories like the transformation and the scattering cancellation theory, which enormously expanded the original concepts and suggest innovative metamaterial-based devices. We hereby use the term "diffusionics" to generalize these remarkable achievements in various energy (e.g., heat) and mass (e.g., particles and plasmas) diffusion systems. For clarity, we categorize the numerous studies appeared during the last decade by diffusion field (i.e., heat, particles, and plasmas) and discuss them from three different perspectives: the theoretical perspective, to detail how the transformation principle is applied to each diffusion field; the application perspective, to introduce various intriguing metamaterial-based devices, such as cloaks and radiative coolers; and the physics perspective, to connect with concepts of recent concern, such as non-Hermitian topology, nonreciprocal transport, and spatiotemporal modulation. We also discuss the possibility of controlling diffusion processes beyond metamaterials. Finally, we point out several future directions for diffusion metamaterial research, including the integration with artificial intelligence and topology concepts. △ Less

Submitted 15 February, 2024; v1 submitted 9 September, 2023; originally announced September 2023.

Comments: This review article has been published in Reviews of Modern Physics, volume 96, 015002 (2024)

Journal ref: Reviews of Modern Physics, volume 96, 015002 (2024)

arXiv:2308.16057 [pdf, other]

Click Metamaterials: Fast Acquisition of Thermal Conductivity and Functionality Diversities

Authors: Chengmeng Wang, Peng Jin, Fubao Yang, Liujun Xu, Jiping Huang

Abstract: Material science is an important foundation of modern society development, covering significant areas like chemosynthesis and metamaterials. Click chemistry provides a simple and efficient paradigm for achieving molecular diversity by incorporating modified building blocks into compounds. In contrast, most metamaterial designs are still case by case due to lacking a fundamental mechanism for achie… ▽ More Material science is an important foundation of modern society development, covering significant areas like chemosynthesis and metamaterials. Click chemistry provides a simple and efficient paradigm for achieving molecular diversity by incorporating modified building blocks into compounds. In contrast, most metamaterial designs are still case by case due to lacking a fundamental mechanism for achieving reconfigurable thermal conductivities, largely hindering design flexibility and functional diversity. Here, we propose a universal concept of click metamaterials for fast realizing various thermal conductivities and functionalities. Tunable hollow-filled unit cells are constructed to mimic the modified building blocks in click chemistry. Different hollow-filled arrays can generate convertible thermal conductivities from isotropy to anisotropy, allowing click metamaterials to exhibit adaptive thermal functionalities. The straightforward structures enable full-parameter regulation and simplify engineering preparation, making click metamaterials a promising candidate for practical use in other diffusion and wave systems. △ Less

Submitted 6 January, 2024; v1 submitted 30 August, 2023; originally announced August 2023.

Comments: Here, click metamaterials have been proposed and swiftly generate variable thermal conductivities and functionalities by using tunable hollow-filled cells akin to the modified building blocks in click chemistry. This breakthrough holds the promise to transform applications in a range of diffusion and wave systems, thereby having a profound impact on the development of materials science

arXiv:2308.11355 [pdf, ps, other]

Machine learning assisted exploration for affine Deligne-Lusztig varieties

Authors: Bin Dong, Xuhua He, Pengfei Jin, Felix Schremmer, Qingchao Yu

Abstract: This paper presents a novel, interdisciplinary study that leverages a Machine Learning (ML) assisted framework to explore the geometry of affine Deligne-Lusztig varieties (ADLV). The primary objective is to investigate the nonemptiness pattern, dimension and enumeration of irreducible components of ADLV. Our proposed framework demonstrates a recursive pipeline of data generation, model training, p… ▽ More This paper presents a novel, interdisciplinary study that leverages a Machine Learning (ML) assisted framework to explore the geometry of affine Deligne-Lusztig varieties (ADLV). The primary objective is to investigate the nonemptiness pattern, dimension and enumeration of irreducible components of ADLV. Our proposed framework demonstrates a recursive pipeline of data generation, model training, pattern analysis, and human examination, presenting an intricate interplay between ML and pure mathematical research. Notably, our data-generation process is nuanced, emphasizing the selection of meaningful subsets and appropriate feature sets. We demonstrate that this framework has a potential to accelerate pure mathematical research, leading to the discovery of new conjectures and promising research directions that could otherwise take significant time to uncover. We rediscover the virtual dimension formula and provide a full mathematical proof of a newly identified problem concerning a certain lower bound of dimension. Furthermore, we extend an open invitation to the readers by providing the source code for computing ADLV and the ML models, promoting further explorations. This paper concludes by sharing valuable experiences and highlighting lessons learned from this collaboration. △ Less

Submitted 22 August, 2023; originally announced August 2023.

Comments: 36 pages

MSC Class: 22E35; 22E67

arXiv:2308.08283 [pdf, other]

CARE: A Large Scale CT Image Dataset and Clinical Applicable Benchmark Model for Rectal Cancer Segmentation

Authors: Hantao Zhang, Weidong Guo, Chenyang Qiu, Shouhong Wan, Bingbing Zou, Wanqin Wang, Peiquan Jin

Abstract: Rectal cancer segmentation of CT image plays a crucial role in timely clinical diagnosis, radiotherapy treatment, and follow-up. Although current segmentation methods have shown promise in delineating cancerous tissues, they still encounter challenges in achieving high segmentation precision. These obstacles arise from the intricate anatomical structures of the rectum and the difficulties in perfo… ▽ More Rectal cancer segmentation of CT image plays a crucial role in timely clinical diagnosis, radiotherapy treatment, and follow-up. Although current segmentation methods have shown promise in delineating cancerous tissues, they still encounter challenges in achieving high segmentation precision. These obstacles arise from the intricate anatomical structures of the rectum and the difficulties in performing differential diagnosis of rectal cancer. Additionally, a major obstacle is the lack of a large-scale, finely annotated CT image dataset for rectal cancer segmentation. To address these issues, this work introduces a novel large scale rectal cancer CT image dataset CARE with pixel-level annotations for both normal and cancerous rectum, which serves as a valuable resource for algorithm research and clinical application development. Moreover, we propose a novel medical cancer lesion segmentation benchmark model named U-SAM. The model is specifically designed to tackle the challenges posed by the intricate anatomical structures of abdominal organs by incorporating prompt information. U-SAM contains three key components: promptable information (e.g., points) to aid in target area localization, a convolution module for capturing low-level lesion details, and skip-connections to preserve and recover spatial information during the encoding-decoding process. To evaluate the effectiveness of U-SAM, we systematically compare its performance with several popular segmentation methods on the CARE dataset. The generalization of the model is further verified on the WORD dataset. Extensive experiments demonstrate that the proposed U-SAM outperforms state-of-the-art methods on these two datasets. These experiments can serve as the baseline for future research and clinical application development. △ Less

Submitted 16 August, 2023; originally announced August 2023.

Comments: 8 pages

arXiv:2308.04020 [pdf, other]

Synthetic Augmentation with Large-scale Unconditional Pre-training

Authors: Jiarong Ye, Haomiao Ni, Peng Jin, Sharon X. Huang, Yuan Xue

Abstract: Deep learning based medical image recognition systems often require a substantial amount of training data with expert annotations, which can be expensive and time-consuming to obtain. Recently, synthetic augmentation techniques have been proposed to mitigate the issue by generating realistic images conditioned on class labels. However, the effectiveness of these methods heavily depends on the repr… ▽ More Deep learning based medical image recognition systems often require a substantial amount of training data with expert annotations, which can be expensive and time-consuming to obtain. Recently, synthetic augmentation techniques have been proposed to mitigate the issue by generating realistic images conditioned on class labels. However, the effectiveness of these methods heavily depends on the representation capability of the trained generative model, which cannot be guaranteed without sufficient labeled training data. To further reduce the dependency on annotated data, we propose a synthetic augmentation method called HistoDiffusion, which can be pre-trained on large-scale unlabeled datasets and later applied to a small-scale labeled dataset for augmented training. In particular, we train a latent diffusion model (LDM) on diverse unlabeled datasets to learn common features and generate realistic images without conditional inputs. Then, we fine-tune the model with classifier guidance in latent space on an unseen labeled dataset so that the model can synthesize images of specific categories. Additionally, we adopt a selective mechanism to only add synthetic samples with high confidence of matching to target labels. We evaluate our proposed method by pre-training on three histopathology datasets and testing on a histopathology dataset of colorectal cancer (CRC) excluded from the pre-training datasets. With HistoDiffusion augmentation, the classification accuracy of a backbone classifier is remarkably improved by 6.4% using a small set of the original labels. Our code is available at https://github.com/karenyyy/HistoDiffAug. △ Less

Submitted 7 August, 2023; originally announced August 2023.

Comments: MICCAI 2023

arXiv:2307.15388 [pdf, other]

An Empirical Study of Large-Scale Data-Driven Full Waveform Inversion

Authors: Peng Jin, Yinan Feng, Shihang Feng, Hanchen Wang, Yinpeng Chen, Benjamin Consolvo, Zicheng Liu, Youzuo Lin

Abstract: This paper investigates the impact of big data on deep learning models to help solve the full waveform inversion (FWI) problem. While it is well known that big data can boost the performance of deep learning models in many tasks, its effectiveness has not been validated for FWI. To address this gap, we present an empirical study that investigates how deep learning models in FWI behave when trained… ▽ More This paper investigates the impact of big data on deep learning models to help solve the full waveform inversion (FWI) problem. While it is well known that big data can boost the performance of deep learning models in many tasks, its effectiveness has not been validated for FWI. To address this gap, we present an empirical study that investigates how deep learning models in FWI behave when trained on OpenFWI, a collection of large-scale, multi-structural, synthetic datasets published recently. In particular, we train and evaluate the FWI models on a combination of 10 2D subsets in OpenFWI that contain 470K pairs of seismic data and velocity maps in total. Our experiments demonstrate that training on the combined dataset yields an average improvement of 13.03% in MAE, 7.19% in MSE and 1.87% in SSIM compared to each split dataset, and an average improvement of 28.60%, 21.55% and 8.22% in the leave-one-out generalization test. We further demonstrate that model capacity needs to scale in accordance with data size for optimal improvement, where our largest model yields an average improvement of 20.06%, 13.39% and 0.72% compared to the smallest one. △ Less

Submitted 24 April, 2024; v1 submitted 28 July, 2023; originally announced July 2023.

arXiv:2307.00458 [pdf]

13.56MHz Rectifying Diodes Based on Metal Halide Perovskite

Authors: Peng Jin, Xuehui Xu, Zeng Chen, Xu Chen, Tianyu Liu, Hanbo Zhu, Xinya Chen, Yang, Yang

Abstract: The increasing use of portable and wireless technologies has led to a growing focus on radio-frequency identification (RFID) tags. Among the various devices in RFID tags, rectifying diodes are the most demanding in terms of high-frequency performance, and these diodes are dominated by organic materials. However, their intrinsic low carrier mobility largely limits the rectifying ability of organic… ▽ More The increasing use of portable and wireless technologies has led to a growing focus on radio-frequency identification (RFID) tags. Among the various devices in RFID tags, rectifying diodes are the most demanding in terms of high-frequency performance, and these diodes are dominated by organic materials. However, their intrinsic low carrier mobility largely limits the rectifying ability of organic diodes. As an alternative, metal halide perovskites (MHPs) possess high carrier mobility, making them potential candidates for higher-frequency applications. Whereas their ion-migration issue may deteriorate their high-frequency performance. In this study, we report rectifying diodes based on MHPs that can rectify an incoming sinusoidal signal at 13.56 MHz. The diodes exhibit a high rectification ratio of 1.9 x 103 at 1 V and can rectify signals at even higher frequencies. We designed a triangular wave detection method to measure the intensity of ion-migration at different frequencies. Interestingly, the ion-migration did not occur at such a high frequency. The high-frequency stagnant ions and excellent carrier mobility make MHPs unexpectedly suitable for high-frequency applications, providing a promising solution to ion-migration issues and paving the way for perovskites in high-frequency areas. △ Less

Submitted 1 July, 2023; originally announced July 2023.

Comments: 19pages, 8 figures, research article, not published

arXiv:2306.15948 [pdf, other]

doi 10.1093/mnras/stad1965

The mHz quasi-regular modulations of 4U 1630--47 during its 1998 outburst

Authors: Qingchang Zhao, Hongxing Yin, Lian Tao, Zixu Yang, Jinlu Qu, Liang Zhang, Shu Zhang, Erlin Qiao, Qingcui Bu, Shujie Zhao, Panping Li, Yiming Huang, Ruican Ma, Ruijing Tang, Pei Jin, Wei Yu, Hexin Liu, Yue Huang, Xiang Ma, Jingyu Xiao, Xuan Zhang, Kang Zhao

Abstract: We present the results of a detailed timing and spectral analysis of the quasi-regular modulation (QRM) phenomenon in the black hole X-ray binary 4U 1630--47 during its 1998 outburst observed by Rossi X-ray Timing Explore (RXTE). We find that the $\sim$ 50-110 mHz QRM is flux dependent, and the QRM is detected with simultaneous low frequency quasi-periodic oscillations (LFQPOs). According to the b… ▽ More We present the results of a detailed timing and spectral analysis of the quasi-regular modulation (QRM) phenomenon in the black hole X-ray binary 4U 1630--47 during its 1998 outburst observed by Rossi X-ray Timing Explore (RXTE). We find that the $\sim$ 50-110 mHz QRM is flux dependent, and the QRM is detected with simultaneous low frequency quasi-periodic oscillations (LFQPOs). According to the behavior of the power density spectrum, we divide the observations into four groups. In the first group, namely behavior A, LFQPOs are detected, but no mHz QRM. The second group, namely behavior B, a QRM with frequency above $\sim$ 88 mHz is detected and the $\sim$ 5 Hz and $\sim$ 7 Hz LFQPOs are almost overlapping. In the third group, namely behavior C, the QRM frequency below $\sim$ 88 mHz is detected and the LFQPOs are significantly separated. In the forth group, namely behavior D, neither QRM nor LFQPOs are detected. We study the energy-dependence of the fractional rms, centroid frequency, and phase-lag of QRM and LFQPOs for behavior B and C. We then study the evolution of QRM and find that the frequency of QRM increases with hardness, while its rms decreases with hardness. We also analyze the spectra of each observation, and find that the QRM rms of behavior B has a positive correlation with $\rm F_{\rm powerlaw}$ / $\rm F_{\rm total}$. Finally, we give our understanding for this mHz QRM phenomena. △ Less

Submitted 28 June, 2023; originally announced June 2023.

Comments: 14pages, 15 figures

arXiv:2306.12456 [pdf, other]

Pushing the Limits of Machine Design: Automated CPU Design with AI

Authors: Shuyao Cheng, Pengwei Jin, Qi Guo, Zidong Du, Rui Zhang, Yunhao Tian, Xing Hu, Yongwei Zhao, Yifan Hao, Xiangtao Guan, Husheng Han, Zhengyue Zhao, Ximing Liu, Ling Li, Xishan Zhang, Yuejie Chu, Weilong Mao, Tianshi Chen, Yunji Chen

Abstract: Design activity -- constructing an artifact description satisfying given goals and constraints -- distinguishes humanity from other animals and traditional machines, and endowing machines with design abilities at the human level or beyond has been a long-term pursuit. Though machines have already demonstrated their abilities in designing new materials, proteins, and computer programs with advanced… ▽ More Design activity -- constructing an artifact description satisfying given goals and constraints -- distinguishes humanity from other animals and traditional machines, and endowing machines with design abilities at the human level or beyond has been a long-term pursuit. Though machines have already demonstrated their abilities in designing new materials, proteins, and computer programs with advanced artificial intelligence (AI) techniques, the search space for designing such objects is relatively small, and thus, "Can machines design like humans?" remains an open question. To explore the boundary of machine design, here we present a new AI approach to automatically design a central processing unit (CPU), the brain of a computer, and one of the world's most intricate devices humanity have ever designed. This approach generates the circuit logic, which is represented by a graph structure called Binary Speculation Diagram (BSD), of the CPU design from only external input-output observations instead of formal program code. During the generation of BSD, Monte Carlo-based expansion and the distance of Boolean functions are used to guarantee accuracy and efficiency, respectively. By efficiently exploring a search space of unprecedented size 10^{10^{540}}, which is the largest one of all machine-designed objects to our best knowledge, and thus pushing the limits of machine design, our approach generates an industrial-scale RISC-V CPU within only 5 hours. The taped-out CPU successfully runs the Linux operating system and performs comparably against the human-designed Intel 80486SX CPU. In addition to learning the world's first CPU only from input-output observations, which may reform the semiconductor industry by significantly reducing the design cycle, our approach even autonomously discovers human knowledge of the von Neumann architecture. △ Less

Submitted 27 June, 2023; v1 submitted 21 June, 2023; originally announced June 2023.

Comments: 28 pages

arXiv:2306.12386 [pdf, other]

$\mathbf{\mathbb{E}^{FWI}}$: Multi-parameter Benchmark Datasets for Elastic Full Waveform Inversion of Geophysical Properties

Authors: Shihang Feng, Hanchen Wang, Chengyuan Deng, Yinan Feng, Yanhua Liu, Min Zhu, Peng Jin, Yinpeng Chen, Youzuo Lin

Abstract: Elastic geophysical properties (such as P- and S-wave velocities) are of great importance to various subsurface applications like CO$_2$ sequestration and energy exploration (e.g., hydrogen and geothermal). Elastic full waveform inversion (FWI) is widely applied for characterizing reservoir properties. In this paper, we introduce $\mathbf{\mathbb{E}^{FWI}}$, a comprehensive benchmark dataset that… ▽ More Elastic geophysical properties (such as P- and S-wave velocities) are of great importance to various subsurface applications like CO$_2$ sequestration and energy exploration (e.g., hydrogen and geothermal). Elastic full waveform inversion (FWI) is widely applied for characterizing reservoir properties. In this paper, we introduce $\mathbf{\mathbb{E}^{FWI}}$, a comprehensive benchmark dataset that is specifically designed for elastic FWI. $\mathbf{\mathbb{E}^{FWI}}$ encompasses 8 distinct datasets that cover diverse subsurface geologic structures (flat, curve, faults, etc). The benchmark results produced by three different deep learning methods are provided. In contrast to our previously presented dataset (pressure recordings) for acoustic FWI (referred to as OpenFWI), the seismic dataset in $\mathbf{\mathbb{E}^{FWI}}$ has both vertical and horizontal components. Moreover, the velocity maps in $\mathbf{\mathbb{E}^{FWI}}$ incorporate both P- and S-wave velocities. While the multicomponent data and the added S-wave velocity make the data more realistic, more challenges are introduced regarding the convergence and computational cost of the inversion. We conduct comprehensive numerical experiments to explore the relationship between P-wave and S-wave velocities in seismic data. The relation between P- and S-wave velocities provides crucial insights into the subsurface properties such as lithology, porosity, fluid content, etc. We anticipate that $\mathbf{\mathbb{E}^{FWI}}$ will facilitate future research on multiparameter inversions and stimulate endeavors in several critical research topics of carbon-zero and new energy exploration. All datasets, codes and relevant information can be accessed through our website at https://efwi-lanl.github.io/ △ Less

Submitted 7 September, 2023; v1 submitted 21 June, 2023; originally announced June 2023.

Comments: 20 pages, 11 figures

arXiv:2306.10750 [pdf, other]

WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation

Authors: Zesen Cheng, Peng Jin, Hao Li, Kehan Li, Siheng Li, Xiangyang Ji, Chang Liu, Jie Chen

Abstract: The top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturbed by Inferior Positive (IP) errors due to the lack of prior object information. Nevertheless, we di… ▽ More The top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturbed by Inferior Positive (IP) errors due to the lack of prior object information. Nevertheless, we discover that two types of methods are highly complementary for restraining respective weaknesses but the direct average combination leads to harmful interference. In this context, we build Win-win Cooperation (WiCo) to exploit complementary nature of two types of methods on both interaction and integration aspects for achieving a win-win improvement. For the interaction aspect, Complementary Feature Interaction (CFI) provides fine-grained information to top-down branch and introduces prior object information to bottom-up branch for complementary feature enhancement. For the integration aspect, Gaussian Scoring Integration (GSI) models the gaussian performance distributions of two branches and weightedly integrates results by sampling confident scores from the distributions. With our WiCo, several prominent top-down and bottom-up combinations achieve remarkable improvements on three common datasets with reasonable extra costs, which justifies effectiveness and generality of our method. △ Less

Submitted 19 June, 2023; originally announced June 2023.

Comments: Accepted to IJCAI2023

arXiv:2306.05445 [pdf, other]

Towards Predicting Equilibrium Distributions for Molecular Systems with Deep Learning

Authors: Shuxin Zheng, Jiyan He, Chang Liu, Yu Shi, Ziheng Lu, Weitao Feng, Fusong Ju, Jiaxi Wang, Jianwei Zhu, Yaosen Min, He Zhang, Shidi Tang, Hongxia Hao, Peiran Jin, Chi Chen, Frank Noé, Haiguang Liu, Tie-Yan Liu

Abstract: Advances in deep learning have greatly improved structure prediction of molecules. However, many macroscopic observations that are important for real-world applications are not functions of a single molecular structure, but rather determined from the equilibrium distribution of structures. Traditional methods for obtaining these distributions, such as molecular dynamics simulation, are computation… ▽ More Advances in deep learning have greatly improved structure prediction of molecules. However, many macroscopic observations that are important for real-world applications are not functions of a single molecular structure, but rather determined from the equilibrium distribution of structures. Traditional methods for obtaining these distributions, such as molecular dynamics simulation, are computationally expensive and often intractable. In this paper, we introduce a novel deep learning framework, called Distributional Graphormer (DiG), in an attempt to predict the equilibrium distribution of molecular systems. Inspired by the annealing process in thermodynamics, DiG employs deep neural networks to transform a simple distribution towards the equilibrium distribution, conditioned on a descriptor of a molecular system, such as a chemical graph or a protein sequence. This framework enables efficient generation of diverse conformations and provides estimations of state densities. We demonstrate the performance of DiG on several molecular tasks, including protein conformation sampling, ligand structure sampling, catalyst-adsorbate sampling, and property-guided structure generation. DiG presents a significant advancement in methodology for statistically understanding molecular systems, opening up new research opportunities in molecular science. △ Less

Submitted 8 June, 2023; originally announced June 2023.

Comments: 80 pages, 11 figures

arXiv:2306.04565 [pdf, ps, other]

Prime sum graphs and the induced trees they contain

Authors: Ernie Croot, Patrick Jin

Abstract: In this paper we show that prime sum graphs on $n$ vertices -- which are graphs on vertex set $\{1,2,...,n\}$ where $ij$ is an edge when $i+j$ is prime -- contain all trees with at most $\exp( c \log n / \log\log n)$ vertices as induced subgraphs. We also prove some results for related graphs, and end with some unsolved problems. In this paper we show that prime sum graphs on $n$ vertices -- which are graphs on vertex set $\{1,2,...,n\}$ where $ij$ is an edge when $i+j$ is prime -- contain all trees with at most $\exp( c \log n / \log\log n)$ vertices as induced subgraphs. We also prove some results for related graphs, and end with some unsolved problems. △ Less

Submitted 7 June, 2023; originally announced June 2023.

Comments: This was part of an undergrad research project that Patrick did with me last year

arXiv:2305.18498 [pdf, other]

ANPL: Towards Natural Programming with Interactive Decomposition

Authors: Di Huang, Ziyuan Nan, Xing Hu, Pengwei Jin, Shaohui Peng, Yuanbo Wen, Rui Zhang, Zidong Du, Qi Guo, Yewen Pu, Yunji Chen

Abstract: Though LLMs are capable of generating plausible programs, it's challenging to interact with the LLMs further to revise the program, especially if the user's specific requirements are different from the initial proposal. In this paper, we introduce ANPL, an interactive programming system that ensures users can always refine the generated code towards their specific programmatic intents via structur… ▽ More Though LLMs are capable of generating plausible programs, it's challenging to interact with the LLMs further to revise the program, especially if the user's specific requirements are different from the initial proposal. In this paper, we introduce ANPL, an interactive programming system that ensures users can always refine the generated code towards their specific programmatic intents via structured decompositions. Borrowing the paradigm of sketching from program synthesis, an ANPL program consists of a set of input-outputs that it must satisfy, a ``sketch'' -- control/data flow expressed in precise code (e.g. Python), and ``holes'' -- sub-modules to be implemented by the LLM specified with natural language. The user revises an ANPL program by either modifying the sketch, changing the language used to describe the holes, or providing additional input-outputs to a particular hole, turning it into a sub-ANPL program that can be solved recursively. This workflow allows the users to offload programming burdens to the LLM as much as possible while retaining the ability to pinpoint and resolve bugs locally, without exposing the rest of the program to the LLM. We deploy ANPL on the Abstraction and Reasoning Corpus (ARC), a set of unique tasks that are challenging for state-of-the-art AI systems, showing it outperforms baseline programming systems that (a) without the ability to decompose tasks interactively and (b) without the guarantee that the modules can be correctly composed together. Additional evaluations on APPS, HumanEval, and real-world programming tasks have validated that the ANPL framework is applicable to multiple programming domains. We release the ANPL solutions to the ARC tasks as a dataset, providing insights into how humans decompose novel tasks programmatically. See our code at https://iprc-dip.github.io/ANPL/. △ Less

Submitted 30 November, 2023; v1 submitted 29 May, 2023; originally announced May 2023.

arXiv:2305.18084 [pdf, other]

Assess and Summarize: Improve Outage Understanding with Large Language Models

Authors: Pengxiang Jin, Shenglin Zhang, Minghua Ma, Haozhe Li, Yu Kang, Liqun Li, Yudong Liu, Bo Qiao, Chaoyun Zhang, Pu Zhao, Shilin He, Federica Sarro, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang

Abstract: Cloud systems have become increasingly popular in recent years due to their flexibility and scalability. Each time cloud computing applications and services hosted on the cloud are affected by a cloud outage, users can experience slow response times, connection issues or total service disruption, resulting in a significant negative business impact. Outages are usually comprised of several concurri… ▽ More Cloud systems have become increasingly popular in recent years due to their flexibility and scalability. Each time cloud computing applications and services hosted on the cloud are affected by a cloud outage, users can experience slow response times, connection issues or total service disruption, resulting in a significant negative business impact. Outages are usually comprised of several concurring events/source causes, and therefore understanding the context of outages is a very challenging yet crucial first step toward mitigating and resolving outages. In current practice, on-call engineers with in-depth domain knowledge, have to manually assess and summarize outages when they happen, which is time-consuming and labor-intensive. In this paper, we first present a large-scale empirical study investigating the way on-call engineers currently deal with cloud outages at Microsoft, and then present and empirically validate a novel approach (dubbed Oasis) to help the engineers in this task. Oasis is able to automatically assess the impact scope of outages as well as to produce human-readable summarization. Specifically, Oasis first assesses the impact scope of an outage by aggregating relevant incidents via multiple techniques. Then, it generates a human-readable summary by leveraging fine-tuned large language models like GPT-3.x. The impact assessment component of Oasis was introduced in Microsoft over three years ago, and it is now widely adopted, while the outage summarization component has been recently introduced, and in this article we present the results of an empirical evaluation we carried out on 18 real-world cloud systems as well as a human-based evaluation with outage owners. The results show that Oasis can effectively and efficiently summarize outages, and lead Microsoft to deploy its first prototype which is currently under experimental adoption by some of the incident teams. △ Less

Submitted 29 May, 2023; originally announced May 2023.

arXiv:2305.13314 [pdf, other]

Auto-Linear Phenomenon in Subsurface Imaging

Authors: Yinan Feng, Yinpeng Chen, Peng Jin, Shihang Feng, Zicheng Liu, Youzuo Lin

Abstract: Subsurface imaging involves solving full waveform inversion (FWI) to predict geophysical properties from measurements. This problem can be reframed as an image-to-image translation, with the usual approach being to train an encoder-decoder network using paired data from two domains: geophysical property and measurement. A recent seminal work (InvLINT) demonstrates there is only a linear mapping be… ▽ More Subsurface imaging involves solving full waveform inversion (FWI) to predict geophysical properties from measurements. This problem can be reframed as an image-to-image translation, with the usual approach being to train an encoder-decoder network using paired data from two domains: geophysical property and measurement. A recent seminal work (InvLINT) demonstrates there is only a linear mapping between the latent spaces of the two domains, and the decoder requires paired data for training. This paper extends this direction by demonstrating that only linear mapping necessitates paired data, while both the encoder and decoder can be learned from their respective domains through self-supervised learning. This unveils an intriguing phenomenon (named Auto-Linear) where the self-learned features of two separate domains are automatically linearly correlated. Compared with existing methods, our Auto-Linear has four advantages: (a) solving both forward and inverse modeling simultaneously, (b) applicable to different subsurface imaging tasks and achieving markedly better results than previous methods, (c)enhanced performance, especially in scenarios with limited paired data and in the presence of noisy data, and (d) strong generalization ability of the trained encoder and decoder. △ Less

Submitted 21 May, 2024; v1 submitted 27 April, 2023; originally announced May 2023.

arXiv:2305.12218 [pdf, other]

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

Authors: Peng Jin, Hao Li, Zesen Cheng, Jinfa Huang, Zhennan Wang, Li Yuan, Chang Liu, Jie Chen

Abstract: Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate t… ▽ More Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate the conceptualizing and reasoning process of human beings. For disentangled conceptualization, we divide the coarse feature into multiple latent factors related to semantic concepts. For set-to-set alignment, where a set of visual concepts correspond to a set of textual concepts, we propose an adaptive pooling method to aggregate semantic concepts to address the partial matching. In particular, since we encode concepts independently in only a few dimensions, DiCoSA is superior at efficiency and granularity, ensuring fine-grained interactions using a similar computational complexity as coarse-grained alignment. Extensive experiments on five datasets, including MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that our method outperforms the existing state-of-the-art methods. △ Less

Submitted 20 May, 2023; originally announced May 2023.

Comments: IJCAI 2023

arXiv:2305.10049 [pdf, other]

TG-VQA: Ternary Game of Video Question Answering

Authors: Hao Li, Peng Jin, Zesen Cheng, Songyang Zhang, Kai Chen, Zhennan Wang, Chang Liu, Jie Chen

Abstract: Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-grained visual-linguistic alignments. In this work, we innovatively resort to game theory, which can s… ▽ More Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-grained visual-linguistic alignments. In this work, we innovatively resort to game theory, which can simulate complicated relationships among multiple players with specific interaction strategies, e.g., video, question, and answer as ternary players, to achieve fine-grained alignment for VideoQA task. Specifically, we carefully design a VideoQA-specific interaction strategy to tailor the characteristics of VideoQA, which can mathematically generate the fine-grained visual-linguistic alignment label without label-intensive efforts. Our TG-VQA outperforms existing state-of-the-art by a large margin (more than 5%) on long-term and short-term VideoQA datasets, verifying its effectiveness and generalization ability. Thanks to the guidance of game-theoretic interaction, our model impressively convergences well on limited data (${10}^4 ~videos$), surpassing most of those pre-trained on large-scale data ($10^7~videos$). △ Less

Submitted 18 May, 2023; v1 submitted 17 May, 2023; originally announced May 2023.

Comments: IJCAI 2023

arXiv:2304.07143 [pdf, other]

Car-Following Models: A Multidisciplinary Review

Authors: Tianya Terry Zhang, Ph. D., Peter J. Jin, Ph. D., Sean T. McQuade, Ph. D., Alexandre Bayen, Ph. D., Benedetto Piccoli

Abstract: Car-following (CF) algorithms are crucial components of traffic simulations and have been integrated into many production vehicles equipped with Advanced Driving Assistance Systems (ADAS). Insights from the model of car-following behavior help us understand the causes of various macro phenomena that arise from interactions between pairs of vehicles. Car-following models encompass multiple discipli… ▽ More Car-following (CF) algorithms are crucial components of traffic simulations and have been integrated into many production vehicles equipped with Advanced Driving Assistance Systems (ADAS). Insights from the model of car-following behavior help us understand the causes of various macro phenomena that arise from interactions between pairs of vehicles. Car-following models encompass multiple disciplines, including traffic engineering, physics, dynamic system control, cognitive science, machine learning, and reinforcement learning. This paper presents an extensive survey that highlights the differences, complementarities, and overlaps among microscopic traffic flow and control models based on their underlying principles and design logic. It reviews representative algorithms, ranging from theory-based kinematic models, Psycho-Physical Models, and Adaptive cruise control models to data-driven algorithms like Reinforcement Learning (RL) and Imitation Learning (IL). The manuscript discusses the strengths and limitations of these models and explores their applications in different contexts. This review synthesizes existing researches across different domains to fill knowledge gaps and offer guidance for future research by identifying the latest trends in car following models and their applications. △ Less

Submitted 5 March, 2024; v1 submitted 14 April, 2023; originally announced April 2023.

Showing 1–50 of 444 results for author: Jin, P