[v4] libcamera: software_isp: Add support for raw monochrome formats
diff mbox series

Message ID 406a6112-8114-4fee-9c73-028009974ae2@laposte.net
State New
Headers show
Series
  • [v4] libcamera: software_isp: Add support for raw monochrome formats
Related show

Commit Message

Alain Cousinie Oct. 8, 2026, 11:20 a.m. UTC
Implement support for monochrome packed and unpacked formats in the software ISP and shaders.

Signed-off-by: Alain Cousinié <alain.cousinie@laposte.net>
---
Hi,

A big thank you to the reviewers for all their insightful feedback, 
which helped me significantly improve the code and fix major issues.
Special thanks to:
    Barnabás Pőcze <barnabas.pocze@ideasonboard.com>
    Milan Zamazal <mzamazal@redhat.com>

Changes since v3:
- Completely refactored Mono histogram routines in swstats_cpu.cpp:
  Implemented a horizontal sampling step of 4 pixels (x += 4) for optimal 
  processing performance, while maintaining 100% vertical line coverage 
  to ensure robust handling of high-contrast industrial targets like 
  book scans and barcodes.
- Streamlined EGL shader selection logic and adjusted black level values 
  to fix the black screen rendering issue on packed formats.
- Enforced single-channel pixel evaluation loops with deferred RGB vector 
  equalization once per frame to eliminate redundant CPU math.
- Fixed code style formatting rules to adhere to standard camelCase naming.
- Structured and optimized GLSL fragment shaders (raw_mono_*.frag):
  * Unified macro naming definitions (MONO_RAW8, MONO_RAW10 and MONO_RAW12) 
    between unpacked and packed shader implementations for clarity.
  * Optimized rendering pipelines by calling applyContrast() only once per 
    pixel on the shared scalar gray value.

*** Call for Testing ***
This patch has been successfully tested on an HP 14-eu0xxx laptop 
with the OmniVision OG0VA1B monochrome sensor (using both CPU and 
GPU/EGL pipelines in qcam). Testing on the og0ve1b module is currently 
in progress.

Since my hardware setup is limited, I would greatly appreciate it if 
anyone with compatible sensors could help test the following specific 
monochrome pipelines and provide feedback:
- Native 8-bit monochrome (R8)
- 10-bit packed monochrome (R10_CSI2P) on other platforms
- 12-bit unpacked and packed monochrome formats (R12 / R12_CSI2P)

Best regards,
Alain

 .../internal/software_isp/swstats_cpu.h       |  11 ++
 src/libcamera/shaders/meson.build             |   2 +
 src/libcamera/shaders/raw_mono_1x_packed.frag |  84 ++++++++++++
 src/libcamera/shaders/raw_mono_unpacked.frag  |  70 ++++++++++
 src/libcamera/software_isp/debayer_cpu.cpp    |  96 ++++++++++++++
 src/libcamera/software_isp/debayer_cpu.h      |   5 +
 src/libcamera/software_isp/debayer_egl.cpp    |  55 ++++++++
 src/libcamera/software_isp/swstats_cpu.cpp    | 120 ++++++++++++++++++
 8 files changed, 443 insertions(+)
 create mode 100644 src/libcamera/shaders/raw_mono_1x_packed.frag
 create mode 100644 src/libcamera/shaders/raw_mono_unpacked.frag

Comments

Barnabás Pőcze Oct. 8, 2026, 3:03 p.m. UTC | #1
Hi

2026. 10. 08. 13:20 keltezéssel, Alain Cousinie írta:
> Implement support for monochrome packed and unpacked formats in the software ISP and shaders.
> 
> Signed-off-by: Alain Cousinié <alain.cousinie@laposte.net>
> ---
> Hi,
> 
> A big thank you to the reviewers for all their insightful feedback,
> which helped me significantly improve the code and fix major issues.
> Special thanks to:
>      Barnabás Pőcze <barnabas.pocze@ideasonboard.com>
>      Milan Zamazal <mzamazal@redhat.com>
> 
> Changes since v3:
> - Completely refactored Mono histogram routines in swstats_cpu.cpp:
>    Implemented a horizontal sampling step of 4 pixels (x += 4) for optimal
>    processing performance, while maintaining 100% vertical line coverage
>    to ensure robust handling of high-contrast industrial targets like
>    book scans and barcodes.
> - Streamlined EGL shader selection logic and adjusted black level values
>    to fix the black screen rendering issue on packed formats.
> - Enforced single-channel pixel evaluation loops with deferred RGB vector
>    equalization once per frame to eliminate redundant CPU math.
> - Fixed code style formatting rules to adhere to standard camelCase naming.
> - Structured and optimized GLSL fragment shaders (raw_mono_*.frag):
>    * Unified macro naming definitions (MONO_RAW8, MONO_RAW10 and MONO_RAW12)
>      between unpacked and packed shader implementations for clarity.
>    * Optimized rendering pipelines by calling applyContrast() only once per
>      pixel on the shared scalar gray value.
> 
> *** Call for Testing ***
> This patch has been successfully tested on an HP 14-eu0xxx laptop
> with the OmniVision OG0VA1B monochrome sensor (using both CPU and
> GPU/EGL pipelines in qcam). Testing on the og0ve1b module is currently
> in progress.
> 
> Since my hardware setup is limited, I would greatly appreciate it if
> anyone with compatible sensors could help test the following specific
> monochrome pipelines and provide feedback:
> - Native 8-bit monochrome (R8)
> - 10-bit packed monochrome (R10_CSI2P) on other platforms
> - 12-bit unpacked and packed monochrome formats (R12 / R12_CSI2P)
> 
> Best regards,
> Alain
> 
>   .../internal/software_isp/swstats_cpu.h       |  11 ++
>   src/libcamera/shaders/meson.build             |   2 +
>   src/libcamera/shaders/raw_mono_1x_packed.frag |  84 ++++++++++++
>   src/libcamera/shaders/raw_mono_unpacked.frag  |  70 ++++++++++
>   src/libcamera/software_isp/debayer_cpu.cpp    |  96 ++++++++++++++
>   src/libcamera/software_isp/debayer_cpu.h      |   5 +
>   src/libcamera/software_isp/debayer_egl.cpp    |  55 ++++++++
>   src/libcamera/software_isp/swstats_cpu.cpp    | 120 ++++++++++++++++++
>   8 files changed, 443 insertions(+)
>   create mode 100644 src/libcamera/shaders/raw_mono_1x_packed.frag
>   create mode 100644 src/libcamera/shaders/raw_mono_unpacked.frag
> 
> [...]
> diff --git a/src/libcamera/software_isp/debayer_cpu.cpp b/src/libcamera/software_isp/debayer_cpu.cpp
> index ce8b3c647..6f6c1739c 100644
> --- a/src/libcamera/software_isp/debayer_cpu.cpp
> +++ b/src/libcamera/software_isp/debayer_cpu.cpp
> @@ -424,6 +424,72 @@ void DebayerCpu::debayer12P_RGRG_BGR888(uint8_t *dst, const uint8_t *src[])
>   	}
>   }
>   
> +template<bool addAlphaByte, bool ccmEnabled>
> +void DebayerCpu::debayer_Mono(uint8_t *dst, const uint8_t *src[])
> +{
> +	if (inputConfig_.bpp == 8) {

Please remove runtime conditions on the bit width. I have previously suggested
having separate functions for each (bit width, packing) pair, just like it is
done for the color formats. Is there an issue with that?


> +		const uint8_t *cur = src[1];
> +
> +		for (int x = 0; x < (int)window_.width;) {
> +			uint16_t pixel_val = *cur++;

Why `uint16_t` ?


> +
> +			STORE_PIXEL(pixel_val, pixel_val, pixel_val)
> +		}
> +		return;
> +	}
> +
> +	const uint8_t *cur = src[1];
> +	const int shift = (inputConfig_.bpp == 10) ? 6 : 2;
> +	for (int x = 0; x < (int)window_.width;) {
> +		uint8_t low_byte = *cur++;
> +		uint8_t high_byte = *cur++;
> +
> +		uint16_t native_val = (high_byte << 8) | low_byte;
> +		uint16_t pixel_val = native_val >> shift;
> +
> +		STORE_PIXEL(pixel_val, pixel_val, pixel_val)
> +	}
> +}
> +
> +template<bool addAlphaByte, bool ccmEnabled>
> +void DebayerCpu::debayerP_Mono(uint8_t *dst, const uint8_t *src[])
> +{
> +	const uint8_t *cur = src[1];
> +
> +	if (inputConfig_.bpp == 10) {
> +		for (int x = 0; x < (int)window_.width;) {
> +			uint16_t g0 = *cur++;
> +			uint16_t g1 = *cur++;
> +			uint16_t g2 = *cur++;
> +			uint16_t g3 = *cur++;
> +			uint8_t lsb = *cur++;
> +
> +			uint16_t gray0 = (g0 << 2) | (lsb & 0x03);
> +			uint16_t gray1 = (g1 << 2) | ((lsb >> 2) & 0x03);
> +			uint16_t gray2 = (g2 << 2) | ((lsb >> 4) & 0x03);
> +			uint16_t gray3 = (g3 << 2) | ((lsb >> 6) & 0x03);
> +
> +			STORE_PIXEL(gray0, gray0, gray0)
> +			STORE_PIXEL(gray1, gray1, gray1)
> +			STORE_PIXEL(gray2, gray2, gray2)
> +			STORE_PIXEL(gray3, gray3, gray3)
> +		}
> +		return;
> +	}
> +	/* 12 bits Packed */
> +	for (int x = 0; x < (int)window_.width;) {
> +		uint16_t g0 = *cur++;
> +		uint16_t g1 = *cur++;
> +		uint8_t lsb = *cur++;
> +
> +		uint16_t gray0 = (g0 << 4) | (lsb  & 0x0F);
> +		uint16_t gray1 = (g1 << 4) | ((lsb >> 4) & 0x0F);
> +
> +		STORE_PIXEL(gray0, gray0, gray0)
> +		STORE_PIXEL(gray1, gray1, gray1)
> +	}
> +}
> +
>   /*
>    * Setup the Debayer object according to the passed in parameters.
>    * Return 0 on success, a negative errno value on failure
> @@ -439,6 +505,21 @@ int DebayerCpu::getInputConfig(PixelFormat inputFormat, DebayerInputConfig &conf
>   						   formats::BGR888,
>   						   formats::XBGR8888,
>   						   formats::ABGR8888 };
> +	if (bayerFormat.order == BayerFormat::Order::MONO) {
> +		if (bayerFormat.packing == BayerFormat::Packing::None) {
> +			config.bpp = (bayerFormat.bitDepth + 7) & ~7;
> +			config.patternSize.width = 2;
> +			config.patternSize.height = 2;
> +		}
> +		else if (bayerFormat.packing == BayerFormat::Packing::CSI2) {
> +			config.bpp = bayerFormat.bitDepth;
> +			config.patternSize.width = (bayerFormat.bitDepth == 10) ? 4 : 2;
> +			config.patternSize.height = 2;
> +		}
> +
> +		config.outputFormats = outputFormats;
> +		return 0;
> +	}
>   
>   	if ((bayerFormat.bitDepth == 8 || bayerFormat.bitDepth == 10 || bayerFormat.bitDepth == 12) &&
>   	    bayerFormat.packing == BayerFormat::Packing::None &&
> @@ -525,6 +606,21 @@ int DebayerCpu::setDebayerFunctions(PixelFormat inputFormat,
>   		return -EINVAL;
>   	};
>   
> +	if (bayerFormat.order == BayerFormat::Order::MONO) {
> +		if (outputFormat == formats::XRGB8888 || outputFormat == formats::ARGB8888 ||
> +		    outputFormat == formats::XBGR8888 || outputFormat == formats::ABGR8888) {
> +			addAlphaByte = true;
> +		}
> +		if (bayerFormat.packing == BayerFormat::Packing::None) {
> +			SET_DEBAYER_METHODS(debayer_Mono, debayer_Mono)
> +			return 0;
> +		} else if (bayerFormat.packing == BayerFormat::Packing::CSI2) {
> +			SET_DEBAYER_METHODS(debayerP_Mono, debayerP_Mono)
> +			return 0;
> +		}
> +		return invalidFmt();
> +	}
> +
>   	switch (outputFormat) {
>   	case formats::XRGB8888:
>   	case formats::ARGB8888:
> [...]
> diff --git a/src/libcamera/software_isp/swstats_cpu.cpp b/src/libcamera/software_isp/swstats_cpu.cpp
> index 7fb77ce7d..a0189e48a 100644
> --- a/src/libcamera/software_isp/swstats_cpu.cpp
> +++ b/src/libcamera/software_isp/swstats_cpu.cpp
> @@ -375,6 +375,71 @@ void SwStatsCpu::statsGBRG12PLine0(const uint8_t *src[], SwIspStats &stats)
>   	SWSTATS_FINISH_LINE_STATS()
>   }
>   
> +void SwStatsCpu::statsMono8Line0(const uint8_t *src[], SwIspStats &stats)
> +{
> +	const uint8_t *cur = src[1] + window_.x;
> +	uint64_t sum = 0;
> +
> +	for (unsigned int x = 0; x < window_.width; x += 4) {
> +		uint8_t val = *cur;
> +		sum += val;
> +		stats.yHistogram[val >> 2] += 4;
> +		cur += 4;
> +	}
> +
> +	stats.sum_.r() += (sum << (8 + sumShift_));

I'm a bit confused here. `sumShift_ == 0` here, but I don't know
why there is a shift in the first place?


> +}
> +
> +void SwStatsCpu::statsMono10Line0(const uint8_t *src[], SwIspStats &stats)
> +{
> +	const uint8_t *cur = src[1] + window_.x * 2 + 1;
> +	uint64_t sum = 0;
> +
> +	for (unsigned int x = 0; x < window_.width; x += 4) {
> +		uint8_t val = *cur;
> +		sum += val;
> +		stats.yHistogram[val >> 2] += 4;
> +		cur += 8;

This will take the least significant byte of the 10 bits, no? Shouldn't it be
something like

   uint16_t *curr = ...;

   for (...) {
     uint16_t val = curr[x];
     stats.yHistogram[val >> 4] += 1;
     sum += val;
   }

I'm also not sure why "4" is used. I think it should have a weight of 1, no?


> +	}
> +
> +	stats.sum_.r() += (sum << (8 + sumShift_));
> +}
> +
> +inline void SwStatsCpu::statsMono12Line0(const uint8_t *src[], SwIspStats &stats)

You can drop the `inline`, it should be in the declaration if you want to make it inline.


> +{
> +	statsMono10Line0(src, stats);
> +}
> +
> +void SwStatsCpu::statsMono10PLine0(const uint8_t *src[], SwIspStats &stats)
> +{
> +	const uint8_t *cur = src[1] + window_.x * 5 / 4;
> +	uint64_t sum = 0;
> +
> +	for (unsigned int x = 0; x < window_.width; x += 4) {
> +		uint8_t val = *cur;
> +		sum += val;
> +		stats.yHistogram[val >> 2] += 4;
> +		cur += 5;
> +	}
> +
> +	stats.sum_.r() += (sum << (8 + sumShift_));
> +}
> +
> +void SwStatsCpu::statsMono12PLine0(const uint8_t *src[], SwIspStats &stats)
> +{
> +	const uint8_t *cur = src[1] + window_.x * 3 / 2;
> +	uint64_t sum = 0;
> +
> +	for (unsigned int x = 0; x < window_.width; x += 4) {
> +		uint8_t val = *cur;
> +		sum += val;
> +		stats.yHistogram[val >> 2] += 4;
> +		cur += 6;
> +	}
> +
> +	stats.sum_.r() += (sum << (8 + sumShift_));
> +}
> +
>   /**
>    * \brief Reset state to start statistics gathering for a new frame
>    * \param[in] frame The frame number
> @@ -494,6 +559,38 @@ int SwStatsCpu::configure(const StreamConfiguration &inputCfg, unsigned int stat
>   
>   	uint8_t bitDepth = bayerFormat.bitDepth;
>   
> +	if (bayerFormat.order == BayerFormat::Order::MONO) {
> +		if (bayerFormat.packing == BayerFormat::Packing::None) {
> +			patternSize_.width = 2;
> +			patternSize_.height = 2;
> +
> +			if (bitDepth == 8) {
> +				stats0_ = &SwStatsCpu::statsMono8Line0;
> +				sumShift_ = 0;
> +			} else if (bitDepth == 10) {
> +				stats0_ = &SwStatsCpu::statsMono10Line0;
> +				sumShift_ = 2;
> +			} else { /* 12 bits */

Why omit the comparison? There is e.g. `formats::R16`. Could you instead
use a `switch` like it is done in other places?


> +				stats0_ = &SwStatsCpu::statsMono12Line0;
> +				sumShift_ = 4;
> +			}
> +		}
> +		else if (bayerFormat.packing == BayerFormat::Packing::CSI2) {
> +			patternSize_.width = (bitDepth == 10) ? 4 : 2;
> +			patternSize_.height = 2;
> +
> +			stats0_ = (bitDepth == 10) ? &SwStatsCpu::statsMono10PLine0 : &SwStatsCpu::statsMono12PLine0;
> +			sumShift_ = 0;

Here's as well, let's check the bit width exactly, and don't assume it's always either 10 or 12.



> +		}
> +
> +		ySkipMask_ = 0x00;
> +		xShift_ = 0;
> +		swapLines_ = false;
> +		processFrame_ = &SwStatsCpu::processMonoFrame2;
> +
> +		return 0;
> +	}
> +
>   	if ((bitDepth == 10 || bitDepth == 12) &&
>   	    bayerFormat.packing == BayerFormat::Packing::CSI2) {
>   		if (bitDepth == 10)
> @@ -594,6 +691,29 @@ void SwStatsCpu::processBayerFrame2(MappedFrameBuffer &in)
>   	}
>   }
>   
> +void SwStatsCpu::processMonoFrame2(MappedFrameBuffer &in)
> +{
> +	const uint8_t *src = in.planes()[0].data();
> +	const uint8_t *linePointers[3] = { nullptr, nullptr, nullptr };
> +
> +	src += window_.y * stride_;
> +
> +
> +	for (unsigned int y = 0; y < window_.height; y++) {
> +		if (y & ySkipMask_) {
> +			src += stride_;
> +			continue;
> +		}
> +
> +		linePointers[1] = src;
> +		(this->*stats0_)(linePointers, stats_[0]);
> +		src += stride_;
> +	}
> +
> +	stats_[0].sum_.g() = stats_[0].sum_.r();
> +	stats_[0].sum_.b() = stats_[0].sum_.r();
> +}
> +
>   /**
>    * \brief Calculate statistics for a frame in one go
>    * \param[in] frame The frame number

Patch
diff mbox series

diff --git a/include/libcamera/internal/software_isp/swstats_cpu.h b/include/libcamera/internal/software_isp/swstats_cpu.h
index 551870921..81a546714 100644
--- a/include/libcamera/internal/software_isp/swstats_cpu.h
+++ b/include/libcamera/internal/software_isp/swstats_cpu.h
@@ -101,8 +101,19 @@  private:
 	/* Bayer 12 bpp packed */
 	void statsBGGR12PLine0(const uint8_t *src[], SwIspStats &stats);
 	void statsGBRG12PLine0(const uint8_t *src[], SwIspStats &stats);
+	/* Mono 8 bpp unpacked */
+	void statsMono8Line0(const uint8_t *src[], SwIspStats &stats);
+	/* Mono 10 bpp unpacked */
+	void statsMono10Line0(const uint8_t *src[], SwIspStats &stats);
+	/* Mono 10 bpp packed */
+	void statsMono10PLine0(const uint8_t *src[], SwIspStats &stats);
+	/* Mono 12 bpp unpacked */
+	void statsMono12Line0(const uint8_t *src[], SwIspStats &stats);
+	/* Mono 12 bpp packed */
+	void statsMono12PLine0(const uint8_t *src[], SwIspStats &stats);
 
 	void processBayerFrame2(MappedFrameBuffer &in);
+	void processMonoFrame2(MappedFrameBuffer &in);
 
 	processFrameFn processFrame_;
 
diff --git a/src/libcamera/shaders/meson.build b/src/libcamera/shaders/meson.build
index c409ff9b0..546053051 100644
--- a/src/libcamera/shaders/meson.build
+++ b/src/libcamera/shaders/meson.build
@@ -7,6 +7,8 @@  shader_files = files([
     'bayer_unpacked.frag',
     'bayer_unpacked.vert',
     'identity.vert',
+    'raw_mono_1x_packed.frag',
+    'raw_mono_unpacked.frag',
 ])
 
 # Generate header from shaders
diff --git a/src/libcamera/shaders/raw_mono_1x_packed.frag b/src/libcamera/shaders/raw_mono_1x_packed.frag
new file mode 100644
index 000000000..e6209ca6e
--- /dev/null
+++ b/src/libcamera/shaders/raw_mono_1x_packed.frag
@@ -0,0 +1,84 @@ 
+/* SPDX-License-Identifier: BSD-2-Clause */
+/*
+ * Copyright (C) 2026, Alain Cousinié
+ * raw_mono_1x_packed.frag - Fragment shader for raw packed monochrome formats (R10P, R12P)
+ */
+
+
+#ifdef GL_ES
+precision highp float;
+#endif
+
+varying vec2            textureOut;
+
+uniform sampler2D       tex_y;
+uniform vec2            tex_size;
+uniform float           gamma;
+uniform float           contrastExp;
+
+uniform vec3 		blacklevel;
+uniform vec3 		awb;
+
+float applyContrast(float value)
+{
+	float isAbove = step(0.5, value);
+	float lowCase = 0.5 * pow(value * 2.0, contrastExp);
+	float highCase = 1.0 - 0.5 * pow((1.0 - value) * 2.0, contrastExp);
+	return mix(lowCase, highCase, isAbove);
+}
+
+float readByte(float pixelX, float pixelY)
+{
+	vec2 uv = vec2(pixelX + 0.5, pixelY + 0.5) / tex_size;
+	return floor(texture2D(tex_y, uv).r * 255.0 + 0.5);
+}
+
+void main(void)
+{
+	vec2 pixelCoords = floor(textureOut * tex_size);
+	float px = pixelCoords.x;
+	float py = pixelCoords.y;
+
+	float grayBits = 0.0;
+	float maxVal = 1.0;
+
+#ifdef MONO_RAW10
+	int index10 = int(mod(px, 4.0));
+	float blockStart10 = floor(px / 4.0) * 5.0;
+
+	float g10 = readByte(blockStart10 + float(index10), py);
+	float lsb10 = readByte(blockStart10 + 4.0, py);
+
+	float divisor10 = exp2(2.0 * float(index10));
+	float lsbBits10 = mod(floor(lsb10 / divisor10), 4.0);
+
+	grayBits = (g10 * 4.0) + lsbBits10;
+	maxVal = 1023.0;
+#endif
+
+#ifdef MONO_RAW12
+	int index12 = int(mod(px, 2.0));
+	float blockStart12 = floor(px / 2.0) * 3.0;
+
+	float g12 = readByte(blockStart12 + float(index12), py);
+	float lsb12 = readByte(blockStart12 + 2.0, py);
+
+	float divisor12 = (index12 == 0) ? 1.0 : 16.0;
+	float lsbBits12 = mod(floor(lsb12 / divisor12), 16.0);
+
+	grayBits = (g12 * 16.0) + lsbBits12;
+	maxVal = 4095.0;
+#endif
+
+	float cVal = grayBits / maxVal;
+
+	cVal = (cVal - blacklevel.g) * awb.g;
+	cVal = clamp(cVal, 0.0, 1.0);
+
+	cVal = applyContrast(cVal);
+
+	vec3 rgb = vec3(cVal);
+	rgb = pow(rgb, vec3(gamma));
+
+	gl_FragColor = vec4(rgb, 1.0);
+}
diff --git a/src/libcamera/shaders/raw_mono_unpacked.frag b/src/libcamera/shaders/raw_mono_unpacked.frag
new file mode 100644
index 000000000..a2c3a832f
--- /dev/null
+++ b/src/libcamera/shaders/raw_mono_unpacked.frag
@@ -0,0 +1,70 @@ 
+/* SPDX-License-Identifier: BSD-2-Clause */
+/*
+ * Copyright (C) 2026, Alain Cousinié
+ * raw_mono_unpacked.frag - Fragment shader for raw unpacked monochrome formats (R8, R10, R12)
+ */
+
+#extension GL_EXT_texture_rg : enable
+
+#ifdef GL_ES
+precision highp float;
+#endif
+
+varying vec4            center;
+varying vec4            xCoord;
+varying vec4            yCoord;
+
+uniform sampler2D 	tex_y;
+uniform float 		gamma;
+uniform float 		contrastExp;
+
+uniform vec3 		blacklevel;
+uniform vec3 		awb;
+
+float applyContrast(float value)
+{
+	float isAbove = step(0.5, value);
+	float lowCase = 0.5 * pow(value * 2.0, contrastExp);
+	float highCase = 1.0 - 0.5 * pow((1.0 - value) * 2.0, contrastExp);
+	return mix(lowCase, highCase, isAbove);
+}
+
+void main(void)
+{
+	vec4 texel = texture2D(tex_y, center.xy);
+
+	float grayBits = 0.0;
+	float maxVal = 1.0;
+
+#ifdef MONO_RAW8
+	grayBits = floor(texel.r * 255.0 + 0.5);
+	maxVal = 255.0;
+#endif
+
+#ifdef MONO_RAW10
+	float byteLow10  = floor(texel.r * 255.0 + 0.5);
+	float byteHigh10 = floor(texel.g * 255.0 + 0.5);
+	grayBits = (byteHigh10 * 256.0) + byteLow10;
+	maxVal = 1023.0;
+#endif
+
+#ifdef MONO_RAW12
+	float byteLow12  = floor(texel.r * 255.0 + 0.5);
+	float byteHigh12 = floor(texel.g * 255.0 + 0.5);
+	float nativeVal12 = (byteHigh12 * 256.0) + byteLow12;
+	grayBits = floor(nativeVal12 * 0.0625);
+	maxVal = 4095.0;
+#endif
+
+	float cVal = grayBits / maxVal;
+
+	cVal = (cVal - blacklevel.g) * awb.g;
+	cVal = clamp(cVal, 0.0, 1.0);
+
+	cVal = applyContrast(cVal);
+
+	vec3 rgb = vec3(cVal);
+	rgb = pow(rgb, vec3(gamma));
+
+	gl_FragColor = vec4(rgb, 1.0);
+}
diff --git a/src/libcamera/software_isp/debayer_cpu.cpp b/src/libcamera/software_isp/debayer_cpu.cpp
index ce8b3c647..6f6c1739c 100644
--- a/src/libcamera/software_isp/debayer_cpu.cpp
+++ b/src/libcamera/software_isp/debayer_cpu.cpp
@@ -424,6 +424,72 @@  void DebayerCpu::debayer12P_RGRG_BGR888(uint8_t *dst, const uint8_t *src[])
 	}
 }
 
+template<bool addAlphaByte, bool ccmEnabled>
+void DebayerCpu::debayer_Mono(uint8_t *dst, const uint8_t *src[])
+{
+	if (inputConfig_.bpp == 8) {
+		const uint8_t *cur = src[1];
+
+		for (int x = 0; x < (int)window_.width;) {
+			uint16_t pixel_val = *cur++;
+
+			STORE_PIXEL(pixel_val, pixel_val, pixel_val)
+		}
+		return;
+	}
+
+	const uint8_t *cur = src[1];
+	const int shift = (inputConfig_.bpp == 10) ? 6 : 2;
+	for (int x = 0; x < (int)window_.width;) {
+		uint8_t low_byte = *cur++;
+		uint8_t high_byte = *cur++;
+
+		uint16_t native_val = (high_byte << 8) | low_byte;
+		uint16_t pixel_val = native_val >> shift;
+
+		STORE_PIXEL(pixel_val, pixel_val, pixel_val)
+	}
+}
+
+template<bool addAlphaByte, bool ccmEnabled>
+void DebayerCpu::debayerP_Mono(uint8_t *dst, const uint8_t *src[])
+{
+	const uint8_t *cur = src[1];
+
+	if (inputConfig_.bpp == 10) {
+		for (int x = 0; x < (int)window_.width;) {
+			uint16_t g0 = *cur++;
+			uint16_t g1 = *cur++;
+			uint16_t g2 = *cur++;
+			uint16_t g3 = *cur++;
+			uint8_t lsb = *cur++;
+
+			uint16_t gray0 = (g0 << 2) | (lsb & 0x03);
+			uint16_t gray1 = (g1 << 2) | ((lsb >> 2) & 0x03);
+			uint16_t gray2 = (g2 << 2) | ((lsb >> 4) & 0x03);
+			uint16_t gray3 = (g3 << 2) | ((lsb >> 6) & 0x03);
+
+			STORE_PIXEL(gray0, gray0, gray0)
+			STORE_PIXEL(gray1, gray1, gray1)
+			STORE_PIXEL(gray2, gray2, gray2)
+			STORE_PIXEL(gray3, gray3, gray3)
+		}
+		return;
+	}
+	/* 12 bits Packed */
+	for (int x = 0; x < (int)window_.width;) {
+		uint16_t g0 = *cur++;
+		uint16_t g1 = *cur++;
+		uint8_t lsb = *cur++;
+
+		uint16_t gray0 = (g0 << 4) | (lsb  & 0x0F);
+		uint16_t gray1 = (g1 << 4) | ((lsb >> 4) & 0x0F);
+
+		STORE_PIXEL(gray0, gray0, gray0)
+		STORE_PIXEL(gray1, gray1, gray1)
+	}
+}
+
 /*
  * Setup the Debayer object according to the passed in parameters.
  * Return 0 on success, a negative errno value on failure
@@ -439,6 +505,21 @@  int DebayerCpu::getInputConfig(PixelFormat inputFormat, DebayerInputConfig &conf
 						   formats::BGR888,
 						   formats::XBGR8888,
 						   formats::ABGR8888 };
+	if (bayerFormat.order == BayerFormat::Order::MONO) {
+		if (bayerFormat.packing == BayerFormat::Packing::None) {
+			config.bpp = (bayerFormat.bitDepth + 7) & ~7;
+			config.patternSize.width = 2;
+			config.patternSize.height = 2;
+		}
+		else if (bayerFormat.packing == BayerFormat::Packing::CSI2) {
+			config.bpp = bayerFormat.bitDepth;
+			config.patternSize.width = (bayerFormat.bitDepth == 10) ? 4 : 2;
+			config.patternSize.height = 2;
+		}
+
+		config.outputFormats = outputFormats;
+		return 0;
+	}
 
 	if ((bayerFormat.bitDepth == 8 || bayerFormat.bitDepth == 10 || bayerFormat.bitDepth == 12) &&
 	    bayerFormat.packing == BayerFormat::Packing::None &&
@@ -525,6 +606,21 @@  int DebayerCpu::setDebayerFunctions(PixelFormat inputFormat,
 		return -EINVAL;
 	};
 
+	if (bayerFormat.order == BayerFormat::Order::MONO) {
+		if (outputFormat == formats::XRGB8888 || outputFormat == formats::ARGB8888 ||
+		    outputFormat == formats::XBGR8888 || outputFormat == formats::ABGR8888) {
+			addAlphaByte = true;
+		}
+		if (bayerFormat.packing == BayerFormat::Packing::None) {
+			SET_DEBAYER_METHODS(debayer_Mono, debayer_Mono)
+			return 0;
+		} else if (bayerFormat.packing == BayerFormat::Packing::CSI2) {
+			SET_DEBAYER_METHODS(debayerP_Mono, debayerP_Mono)
+			return 0;
+		}
+		return invalidFmt();
+	}
+
 	switch (outputFormat) {
 	case formats::XRGB8888:
 	case formats::ARGB8888:
diff --git a/src/libcamera/software_isp/debayer_cpu.h b/src/libcamera/software_isp/debayer_cpu.h
index 2c88c9e1a..909798fa4 100644
--- a/src/libcamera/software_isp/debayer_cpu.h
+++ b/src/libcamera/software_isp/debayer_cpu.h
@@ -119,6 +119,11 @@  private:
 	void debayer12P_GBGB_BGR888(uint8_t *dst, const uint8_t *src[]);
 	template<bool addAlphaByte, bool ccmEnabled>
 	void debayer12P_RGRG_BGR888(uint8_t *dst, const uint8_t *src[]);
+	/* CSI-2 packed and unpacked raw mono */
+	template<bool addAlphaByte, bool ccmEnabled>
+	void debayer_Mono(uint8_t *dst, const uint8_t *src[]);
+	template<bool addAlphaByte, bool ccmEnabled>
+	void debayerP_Mono(uint8_t *dst, const uint8_t *src[]);
 
 	static int getInputConfig(PixelFormat inputFormat, DebayerInputConfig &config);
 	int setupStandardBayerOrder(BayerFormat::Order order);
diff --git a/src/libcamera/software_isp/debayer_egl.cpp b/src/libcamera/software_isp/debayer_egl.cpp
index 300822f9d..23499bce3 100644
--- a/src/libcamera/software_isp/debayer_egl.cpp
+++ b/src/libcamera/software_isp/debayer_egl.cpp
@@ -61,6 +61,17 @@  int DebayerEGL::getInputConfig(PixelFormat inputFormat, DebayerInputConfig &conf
 						   formats::XBGR8888,
 						   formats::ABGR8888 };
 
+	/* Handle Monochrome case using bayerFormat */
+	if (bayerFormat.order == BayerFormat::Order::MONO) {
+		bool isPacked = (bayerFormat.packing == BayerFormat::Packing::CSI2);
+
+		config.bpp = isPacked ? bayerFormat.bitDepth : ((bayerFormat.bitDepth + 7) & ~7);
+		config.patternSize.width = isPacked ? 4 : 2;
+		config.patternSize.height = 2;
+		config.outputFormats = outputFormats;
+		return 0;
+	}
+
 	if ((bayerFormat.bitDepth == 8 || bayerFormat.bitDepth == 10) &&
 	    bayerFormat.packing == BayerFormat::Packing::None &&
 	    isStandardBayerOrder(bayerFormat.order)) {
@@ -166,6 +177,13 @@  int DebayerEGL::initBayerShaders(PixelFormat inputFormat, PixelFormat outputForm
 	shaderStridePixels_ = inputConfig_.stride;
 
 	switch (inputFormat) {
+	/* Monochrome cases: No color: phase, proper initialization to 0 for Rx and Rx_CSI2P */
+	case libcamera::formats::R8:
+	case libcamera::formats::R10:
+	case libcamera::formats::R12:
+		firstRed_x_ = 0.0;
+		firstRed_y_ = 0.0;
+		break;
 	case libcamera::formats::SBGGR8:
 	case libcamera::formats::SBGGR10_CSI2P:
 	case libcamera::formats::SBGGR12_CSI2P:
@@ -197,6 +215,36 @@  int DebayerEGL::initBayerShaders(PixelFormat inputFormat, PixelFormat outputForm
 
 	/* Shader selection */
 	switch (inputFormat) {
+	case libcamera::formats::R8: {
+		fragmentShaderData = raw_mono_unpacked_frag;
+		vertexShaderData = bayer_unpacked_vert;
+		glFormat_ = GL_LUMINANCE;
+		bytesPerPixel_ = 1;
+		egl_.pushEnv(shaderEnv, "#define MONO_RAW8");
+		break;
+	}
+	case libcamera::formats::R10:
+	case libcamera::formats::R12: {
+		BayerFormat bayerFormat = BayerFormat::fromPixelFormat(inputFormat);
+
+		if (bayerFormat.packing == BayerFormat::Packing::None) {
+			fragmentShaderData = raw_mono_unpacked_frag;
+			vertexShaderData = bayer_unpacked_vert;
+			glFormat_ = GL_RG;
+			bytesPerPixel_ = 2;
+		} else {
+			fragmentShaderData = raw_mono_1x_packed_frag;
+			vertexShaderData = identity_vert;
+			glFormat_ = GL_LUMINANCE;
+			bytesPerPixel_ = 1;
+			shaderStridePixels_ = inputConfig_.stride;
+		}
+		if (inputFormat == libcamera::formats::R10)
+			egl_.pushEnv(shaderEnv, "#define MONO_RAW10");
+		else
+			egl_.pushEnv(shaderEnv, "#define MONO_RAW12");
+		break;
+	}
 	case libcamera::formats::SBGGR8:
 	case libcamera::formats::SGBRG8:
 	case libcamera::formats::SGRBG8:
@@ -516,6 +564,13 @@  eGLImage *DebayerEGL::getCachedInputFrameBuffer(FrameBuffer *input, std::optiona
 		eglImageInCache_.emplace_back(fd, std::make_unique<eGLImage>(glFormat_, inputConfig_.stride / bytesPerPixel_, height_, inputConfig_.stride, GL_TEXTURE0, 0));
 		eglImageIn = eglImageInCache_.back().second.get();
 
+		/* Force GL_NEAREST to preserve LSBs of packed monochrome format */
+		BayerFormat bayerFormat = BayerFormat::fromPixelFormat(inputPixelFormat_);
+		if (bayerFormat.order == BayerFormat::Order::MONO && bayerFormat.packing == BayerFormat::Packing::CSI2) {
+			glTexParameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_NEAREST);
+			glTexParameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_NEAREST);
+		}
+
 		if (egl_.createInputDMABufTexture2D(*eglImageIn, input->planes()[0].fd.get()) == 0)
 			return eglImageIn;
 
diff --git a/src/libcamera/software_isp/swstats_cpu.cpp b/src/libcamera/software_isp/swstats_cpu.cpp
index 7fb77ce7d..a0189e48a 100644
--- a/src/libcamera/software_isp/swstats_cpu.cpp
+++ b/src/libcamera/software_isp/swstats_cpu.cpp
@@ -375,6 +375,71 @@  void SwStatsCpu::statsGBRG12PLine0(const uint8_t *src[], SwIspStats &stats)
 	SWSTATS_FINISH_LINE_STATS()
 }
 
+void SwStatsCpu::statsMono8Line0(const uint8_t *src[], SwIspStats &stats)
+{
+	const uint8_t *cur = src[1] + window_.x;
+	uint64_t sum = 0;
+
+	for (unsigned int x = 0; x < window_.width; x += 4) {
+		uint8_t val = *cur;
+		sum += val;
+		stats.yHistogram[val >> 2] += 4;
+		cur += 4;
+	}
+
+	stats.sum_.r() += (sum << (8 + sumShift_));
+}
+
+void SwStatsCpu::statsMono10Line0(const uint8_t *src[], SwIspStats &stats)
+{
+	const uint8_t *cur = src[1] + window_.x * 2 + 1;
+	uint64_t sum = 0;
+
+	for (unsigned int x = 0; x < window_.width; x += 4) {
+		uint8_t val = *cur;
+		sum += val;
+		stats.yHistogram[val >> 2] += 4;
+		cur += 8;
+	}
+
+	stats.sum_.r() += (sum << (8 + sumShift_));
+}
+
+inline void SwStatsCpu::statsMono12Line0(const uint8_t *src[], SwIspStats &stats)
+{
+	statsMono10Line0(src, stats);
+}
+
+void SwStatsCpu::statsMono10PLine0(const uint8_t *src[], SwIspStats &stats)
+{
+	const uint8_t *cur = src[1] + window_.x * 5 / 4;
+	uint64_t sum = 0;
+
+	for (unsigned int x = 0; x < window_.width; x += 4) {
+		uint8_t val = *cur;
+		sum += val;
+		stats.yHistogram[val >> 2] += 4;
+		cur += 5;
+	}
+
+	stats.sum_.r() += (sum << (8 + sumShift_));
+}
+
+void SwStatsCpu::statsMono12PLine0(const uint8_t *src[], SwIspStats &stats)
+{
+	const uint8_t *cur = src[1] + window_.x * 3 / 2;
+	uint64_t sum = 0;
+
+	for (unsigned int x = 0; x < window_.width; x += 4) {
+		uint8_t val = *cur;
+		sum += val;
+		stats.yHistogram[val >> 2] += 4;
+		cur += 6;
+	}
+
+	stats.sum_.r() += (sum << (8 + sumShift_));
+}
+
 /**
  * \brief Reset state to start statistics gathering for a new frame
  * \param[in] frame The frame number
@@ -494,6 +559,38 @@  int SwStatsCpu::configure(const StreamConfiguration &inputCfg, unsigned int stat
 
 	uint8_t bitDepth = bayerFormat.bitDepth;
 
+	if (bayerFormat.order == BayerFormat::Order::MONO) {
+		if (bayerFormat.packing == BayerFormat::Packing::None) {
+			patternSize_.width = 2;
+			patternSize_.height = 2;
+
+			if (bitDepth == 8) {
+				stats0_ = &SwStatsCpu::statsMono8Line0;
+				sumShift_ = 0;
+			} else if (bitDepth == 10) {
+				stats0_ = &SwStatsCpu::statsMono10Line0;
+				sumShift_ = 2;
+			} else { /* 12 bits */
+				stats0_ = &SwStatsCpu::statsMono12Line0;
+				sumShift_ = 4;
+			}
+		}
+		else if (bayerFormat.packing == BayerFormat::Packing::CSI2) {
+			patternSize_.width = (bitDepth == 10) ? 4 : 2;
+			patternSize_.height = 2;
+
+			stats0_ = (bitDepth == 10) ? &SwStatsCpu::statsMono10PLine0 : &SwStatsCpu::statsMono12PLine0;
+			sumShift_ = 0;
+		}
+
+		ySkipMask_ = 0x00;
+		xShift_ = 0;
+		swapLines_ = false;
+		processFrame_ = &SwStatsCpu::processMonoFrame2;
+
+		return 0;
+	}
+
 	if ((bitDepth == 10 || bitDepth == 12) &&
 	    bayerFormat.packing == BayerFormat::Packing::CSI2) {
 		if (bitDepth == 10)
@@ -594,6 +691,29 @@  void SwStatsCpu::processBayerFrame2(MappedFrameBuffer &in)
 	}
 }
 
+void SwStatsCpu::processMonoFrame2(MappedFrameBuffer &in)
+{
+	const uint8_t *src = in.planes()[0].data();
+	const uint8_t *linePointers[3] = { nullptr, nullptr, nullptr };
+
+	src += window_.y * stride_;
+
+
+	for (unsigned int y = 0; y < window_.height; y++) {
+		if (y & ySkipMask_) {
+			src += stride_;
+			continue;
+		}
+
+		linePointers[1] = src;
+		(this->*stats0_)(linePointers, stats_[0]);
+		src += stride_;
+	}
+
+	stats_[0].sum_.g() = stats_[0].sum_.r();
+	stats_[0].sum_.b() = stats_[0].sum_.r();
+}
+
 /**
  * \brief Calculate statistics for a frame in one go
  * \param[in] frame The frame number