A hybrid vision-language artistic images aesthetic evaluation framework based on cross-modal attention and differential evolution
Aesthetic evaluation of artistic images remains a challenging prediction problem, constrained by the subjective nature of human aesthetic cognition and the complex interplay between low-level visual structures and high-level semantic characteristics. In terms of model construction, the major bottleneck is to establish discriminative representations that can sufficiently model nonlinear dependencies across diverse heterogeneous modalities. In this paper, we propose VPT-IAA, a unified vision–language framework that formulates artistic aesthetic assessment as a cross-modal representation learning problem with explicit interaction modeling. The visual modality is encoded through a permutation-based feature transformation that enables structured aggregation of spatial information across multiple dimensions, while structured multi-dimensional aesthetic attribute descriptions are automatically generated by a MiniCPM-V multimodal large language model pipeline that decomposes artistic appreciation into seven independent dimensions across four hierarchical perception layers, and these descriptions are subsequently embedded using a task-oriented Transformer encoder to obtain attribute-consistent textual representations. To model cross-modal dependencies, we introduce a three-pathway attention mechanism that establishes complementary query–key–value interactions between visual and textual features, yielding a coupled representation space. In addition, bilinear pooling is employed to characterize second-order correlations between modalities, allowing the model to capture higher-order aesthetic relationships. Model hyperparameters are optimized via differential evolution to enhance stability and robustness. Experimental evaluation on the LAPIS dataset demonstrates that the proposed formulation achieves accurate and consistent aesthetic prediction, attaining an MAE of 1.796 and PC of 0.9827–representing a 15.8% reduction in MAE relative to the second-best method AesExpert and a 68.6% reduction relative to CNN-based baselines. Further analyses indicate that explicit cross-modal interaction modeling plays a dominant role in performance improvement, and the proposed framework maintains stable behavior across a wide range of artistic styles, including both representational and abstract artworks.
More from Cuppa
Google's new 'CC' is an AI agent that helps families run their householdsGoogle has relaunched its CC AI agent as a household organiser for families, helping manage schedules, school emails, meal plans and shared to-do lists.
SpaceX to fly more NASA crews to space station under expanded $5.92 billion dealNASA has handed SpaceX a $946 million contract for three more astronaut missions to the International Space Station, extending the partnership through 2030.
Coffee on the nanoscale: Graphene oxide membrane removes half the caffeine while retaining key compoundsAustralian researchers at UNSW have developed a graphene oxide membrane that removes about half the caffeine from already-brewed coffee while keeping its flavour compounds intact.
China extending sea dominion to Batanes?China's Coast Guard is now patrolling waters east of the Philippines' Batanes islands, raising fears Beijing is expanding its maritime claims beyond the South China Sea.
Women and girls pay high price in DR Congo Ebola outbreakWomen and girls make up more than half of confirmed Ebola cases in the DR Congo outbreak, largely because they bear the burden of caring for the sick at home.
Macklemore ticket prices rise amid Ed Sheeran tour falloutMacklemore was dropped from Ed Sheeran's tour after making pro-Palestine comments onstage, and has since announced he will donate his $1 million in tour earnings to Palestinian aid.
UK publisher drops books by David WalliamsHarperCollins UK is dropping David Walliams' entire back catalogue after allegations of harassing female staff, going further than its earlier decision to stop new titles.
Uber ordered to pay $40m to family of woman killed after driver left her on highwayA US jury ordered Uber to pay $40 million to the family of a 23-year-old killed after her driver dumped her and a sick friend on a California highway.