Color matching algorithm and color transfer for film and photography

I think these are interesting ideas. Currently my idea is that additive and subtractive are both projections of underlying stimulus ensembles. The goal then is not to model them separately, but to understand the invariant structure that makes them perceptually equivalent, similar or different.
I agree. I came to the same conclusion because either we have an infinite number of color models or we have to create some sort of sign-post similar to cie but for cmy, essentially a subtractive observer.
 
I have argued on a number of occasions for the idea that color is a state variable derived from stimuli ensembles. A lot of emphasis in color science is put on describing perceived color at a local level, whether it be color patches or through some form of segmentation of forms, but few papers focus on the macro-look from the point of view of phenomological modelling based on first principles. Many concepts overemphasize:

- local descriptors
- tristimulus mappings
- patch-based experiments

And underrepresent:

- ensemble structure
- relational invariants
- macroscopic perceptual states

Within this framework the largest ensemble defined by Plook represents the macro-look. The macro-look is an interesting macroscopic property associated with color, that does not describe the color of individual pixels, regions or even objects, but the relationship between colors. We perceive this relationship even when the semantic subject matter changes. It highlights the importance of not only considering the microscopic, but also the mesoscopic and macroscopic scales in perception, which is where the analogy of statistical thermodynamics becomes relevant. Below is an example of transferring the macro-look of a vintage photograph of an Amsterdam tram to a modern photograph of a vintage tram in Amsteram:

Source:

hQgsIND.jpeg


Reference:

Tfiy3lR.jpeg


Color matching result:

pwc1Ib2.jpeg


Images used for research and educational purposes.
 
Last edited:
A few posts back I shared my PhD thesis, which has nothing to do with color science, but has some relevance to my way of approaching the color science domain. For those that understandably have zero interest in reading a twenty year old text on polymer physics, I will briefly explain where the relevance is to my current work. Ultimately it is all about levels of description. Consider this image:

2WT3DhD.png


The image was taken from the 2020 scientific paper Multiscale Modeling of Sickle Cell Anemia (https://link.springer.com/rwe/10.1007/978-3-319-44680-6_67). The image shows an example of mesoscale modelling of complex molecules, which are represented as a sphere, or multiple spheres. The soft sphere is a graphical representation of an interaction potential, where the ensemble interactions between two or more molecules (or fragments of molecules) are replaced by a statistical average interaction. These statistical blobs allow for modelling large scale structures without including all the molecular detail. Yet despite the simpler nature of the blob model, macroscopic thermodynamic properties generally align very well with experimental results. Modelling a complex interacting molecules as soft spheres or blobs is obviously computationally much more efficient. However, there are other reasons for simplifying a model, and two important ones are for tractability and understanding. These abstractions are not arbitrary. In order to model a physical phenomenon at a microscopic, mesoscopic or macroscopic scale, which details can be discarded and which details cannot? Which details can explain a certain phenomenon, and which details are less relevant or even irrelevant?

A second important concept I would like to highlight is the concept of regimes. Regimes are all about scope: under what constraints and conditions is a model expected to be useful and/or accurate? Different regimes generally show different modes of behavior. I have argued in past posts that my approach is a mean-field model applicable to a specific regime: high entropy images with a wide variety of stimuli, where ensemble statistics dominate. I have likened this regime to a fluid with many interactions, where many traditional models in color science in my view are more akin to a gaseous state with limited interactions. Following through on this analogy I thus interpret models such a CIELAB or Oklab as simple equations of state. They approximately describe the state of the perceptual system in a very specific regime (a single stimulus patch against a uniform field), and cannot be applied universally, nor were they ever meant to, even though many people still use them outside of their intended scope.

Lastly, I would like to return to two sets of images that best encapsulate my approach to some aspects of color science, specifically chromatic adaptation, color constancy, the macro-look, and look-transfer. The first image highlights the concept of corresponding states in the mean-field regime, where chromatic adaption is modelled as a perceptual phase separation: a statistical segmentation of a veil layer and a neutral anchor (the constant in color constancy), where color is represented by a higher dimensional state variable, which can for the purposes of more traditional color matching experiments be collapsed to a lower dimensional set of tri-stimulus values:

s04Z0Zi.jpeg


The macro-look and look-transfer are closely tied to the same concepts that drive chromatic adaptation using the same invariant statistical feature, the relationship structure across the ensemble represented by Plook, which allows for the possibility to place a different image in the same set of corresponding states:

M2SOJUz.jpeg


To summarize in my approach color is a property of ensembles under constraints. It is is not fundamentally microscopic, but a state variable emerging from interactions across scale.

Images used for research and educational purposes.
 
Last edited:
This post has two parts, where the second part is likely most exciting for those that want to know where this all is heading. However, first I would like to share some R&D results. As I explained in the past the look-transfer algorithm has been developed under the assumption that the stimuli in the source and reference frame are identical or at least highly similar. If not the method will likely fail to give satisfactory results. This is a particular challenge for closeups, where dissimilar elements have the greatest impact. Under such circumstances a colorist will focus on specific similar elements in the source and reference, where as many a colorist will attest to skin tones are of particular importance. So, we asked ourselves: what if we replicate this process? How much of the grade is anchored to the skin tones? As it turns out quite a bit in many cases. To test this we cropped the source and reference frames to just include the faces. Here we would like to share a nice example:

Source:

kN4cNcw.jpeg


Reference:

TRxh9uP.jpeg


Color matching results:

QCXXxCe.jpeg


The example highlights the robustness of the algorithm as it extrapolates adjustments to the skin tones to the rest of the frame, but equally if not more important from a color science perspective is the observation that the overall look and feel of the grade is largely encapsulated by the skin tones for many examples. It also highlights the relevance of Plook x Plum as an invariant at a mesoscopic level with a high correlation to the macro-look at the macroscopic level.

Now for the second part of this post. We'll very soon be making a few significant announcements, that we are very excited about. We're not quite ready to share the news, but we do want to share this teaser:

fbU2MIq.jpeg


Watch this space for the upcoming announcements!

Images used for research and educational purposes.
 
Last edited:
What relational invariants are you assuming for P_look?

My understanding of your approach is that you are representing the tristimulus color distribution as the product of three terms (P_look P_lum, P_cont). I think my mental model for this is that P_look is a pushforward transport of the distribution P_lum and P_cont that together might be seen as the product of something like albedo and shading from traditional intrinsic image decomposition. Factoring away P_look from an image would make it look neutral, but still shaded (perhaps conceptually a bit like whitebalancing?). In my current mental model of P_look it's ultimately a function that maps tristimulus color to tristimulus color - representable by a LUT or a high dimensional polynomial, etc. Fitting P_look requires factoring P_lum and P_cont from image so that you can approximate f(p_lum * p_cont = color) => color. The difficulty is in what assumptions do you make to factor p_lum and p_cont from the degenerate p_image. This might be the wrong framing though.
 
What relational invariants are you assuming for P_look?

My understanding of your approach is that you are representing the tristimulus color distribution as the product of three terms (P_look P_lum, P_cont). I think my mental model for this is that P_look is a pushforward transport of the distribution P_lum and P_cont that together might be seen as the product of something like albedo and shading from traditional intrinsic image decomposition. Factoring away P_look from an image would make it look neutral, but still shaded (perhaps conceptually a bit like whitebalancing?). In my current mental model of P_look it's ultimately a function that maps tristimulus color to tristimulus color - representable by a LUT or a high dimensional polynomial, etc. Fitting P_look requires factoring P_lum and P_cont from image so that you can approximate f(p_lum * p_cont = color) => color. The difficulty is in what assumptions do you make to factor p_lum and p_cont from the degenerate p_image. This might be the wrong framing though.

Thanks for your reply Robert! I think you're pretty close. In a general sense a scale dependent stimuli stimuli distribution can be determined at any location in the image. Let's denote this Pimage(x,y,,lambda), where x and y denotr the location, and lambda is the scale parameter. The framework subsequently postulates that Pimage(...) can be statistically decomposed as:

Pimage(x,y,lambda) = Plook x Plum(x,y,lamda) x Pcomp(x,y,lambda)

Here Plook is a macroscopic stimuli distribution that defines the macro-look. Plum(x,y,lambda) is a mesoscopic to macroscopic luminance distribution. Finally, Pcomp(x,y,lambda) is a mesoscopic to macroscopic compositional stimuli distribution. Factoring away Plook results in a balanced image, but Plook is generally not uniform, so factoring it away will result in a poorly formed image.
 
An interesting look transfer example, where a vintage look is transferred, where percepts like color balance, contrast, and saturation are ensemble properties captured by Plook x Plum:

Source:

O7bQ3Yb.jpeg


Reference:

OF0qVXc.jpeg


Color matching result:

FUXO3Bw.png


Images used for research and educational purposes.
 
I have been working on a mean-field model for predicting lightness, hue shifts and saturation as a function of the contextual distribution of lightness including the gamut expansion effect. This model is structurally similar to the tristimulus projection for predicting appearance under a color cast. Here are three example sets, where for each set the center square on the left is represented by a constant RGB triplet. The right is the predicted percept triplet against a mid-grey field:

Set 1:

upt9C0Y.png


iYNJA1r.png


2JJkv8V.png


XbxoYN0.png


9CvKTti.png


Set 2:

7uZm0HP.png


y3PJqSj.png


jQfa4Xi.png


z6bmrc6.png


OHvQf9h.png


Set 3:

VKgf4LY.png


wUa8l6K.png


pVY7Ggn.png


xXcoUYX.png


HgPi765.png
 
Last edited:
I should note that the uniform field case may just be considered a constrained variant of the above presented results. I opted for a variegated uniformly distributed field, because it can be considered a toy model for natural images, and allows for including the gamut expansion effect, and perceptual black and white.

Next up are natural images, and more complex geometries that are within the scope of a mean-field assumption:

97932.jpg
 
Last edited:
Upon reflection the lightness prediction model gives reasonable estimates, but there is a systematic bias, where it overestimates the simultaneous contrast effect in both directions. There's also some issues at the boundary conditions. Consequently, it is not at the same level of accuracy as previous models in this framework, so it is back to the drawing board for me.
 
Last edited:
Back
Top