output image resolution details issues #126

bjohn22 · 2025-02-02T15:54:27Z

Any thoughts on why I am not able to reproduce the same resolution reported out there:
Weights: I downloaded weights from Huggingface to local and loaded it from local directory
Model: Janus-Pro-7B

Prompt python "/home/mytemp/Documents/ml_projects/Janus/janus_run_script.py" generate --prompt "Elephant on Dallas Texas street" --seed 142 --guidance 5.0

Here is my text to image snippet:

def generate(input_ids,
width,
height,
temperature: float = 1,
parallel_size: int = 16,
cfg_weight: float = 5,
image_token_num_per_image: int = 576,
patch_size: int = 16):
"""
Internal method that generates image patches from the text input.
"""
torch.cuda.empty_cache()
tokens = torch.zeros((parallel_size * 2, len(input_ids)), dtype=torch.int).to(cuda_device)

# Prepare token embeddings
for i in range(parallel_size * 2):
    tokens[i, :] = input_ids
    # Odd rows get "unconditional" tokens
    if i % 2 != 0:
        tokens[i, 1:-1] = vl_chat_processor.pad_id

inputs_embeds = vl_gpt.language_model.get_input_embeddings()(tokens)
generated_tokens = torch.zeros((parallel_size, image_token_num_per_image), dtype=torch.int).to(cuda_device)

pkv = None
# Loop to generate tokens for image
for i in range(image_token_num_per_image):
    outputs = vl_gpt.language_model.model(inputs_embeds=inputs_embeds, use_cache=True, past_key_values=pkv)
    pkv = outputs.past_key_values
    hidden_states = outputs.last_hidden_state
    logits = vl_gpt.gen_head(hidden_states[:, -1, :])

    # Extract conditional/unconditional logits
    logit_cond = logits[0::2, :]
    logit_uncond = logits[1::2, :]
    logits = logit_uncond + cfg_weight * (logit_cond - logit_uncond)

    probs = torch.softmax(logits / temperature, dim=-1)
    next_token = torch.multinomial(probs, num_samples=1)
    generated_tokens[:, i] = next_token.squeeze(dim=-1)

    # Expand the next_token for conditional/unconditional
    next_token = torch.cat([next_token.unsqueeze(dim=1), next_token.unsqueeze(dim=1)], dim=1).view(-1)
    img_embeds = vl_gpt.prepare_gen_img_embeds(next_token)
    inputs_embeds = img_embeds.unsqueeze(dim=1)

# Decode tokens into patches
patches = vl_gpt.gen_vision_model.decode_code(
    generated_tokens.to(dtype=torch.int),
    shape=[parallel_size, 8, width // patch_size, height // patch_size]
)
return generated_tokens.to(dtype=torch.int), patches

def unpack(dec, width, height, parallel_size=16):
"""
Convert raw patches into final numpy images.
"""
dec = dec.to(torch.float32).cpu().numpy().transpose(0, 2, 3, 1)
dec = np.clip((dec + 1) / 2 * 255, 0, 255)

visual_img = np.zeros((parallel_size, width, height, 3), dtype=np.uint8)
visual_img[:, :, :] = dec
return visual_img

@torch.inference_mode()
def generate_image(prompt, seed, guidance):
"""
Generate multiple images for a given prompt.
"""
torch.cuda.empty_cache()
seed = seed if seed is not None else 12345
torch.manual_seed(seed)
torch.cuda.manual_seed(seed)
np.random.seed(seed)

width = 384
height = 384
parallel_size = 16

with torch.no_grad():
    # Prepare text
    messages = [
        {'role': 'User', 'content': prompt},
        {'role': 'Assistant', 'content': ''}
    ]
    text = vl_chat_processor.apply_sft_template_for_multi_turn_prompts(
        conversations=messages,
        sft_format=vl_chat_processor.sft_format,
        system_prompt=''
    )
    text = text + vl_chat_processor.image_start_tag

    input_ids = torch.LongTensor(tokenizer.encode(text)).to(cuda_device)

    _, patches = generate(
        input_ids,
        width // 16 * 16,
        height // 16 * 16,
        cfg_weight=guidance,
        parallel_size=parallel_size
    )
    images = unpack(patches, width // 16 * 16, height // 16 * 16)

    # Convert to PIL and upscale
    return [
        Image.fromarray(images[i]).resize((1024, 1024), Image.LANCZOS)
        for i in range(parallel_size)
    ]

The text was updated successfully, but these errors were encountered:

bjohn22 · 2025-02-02T15:59:13Z

Here are my files list

bjohn22 changed the title ~~resolution issues~~ output image resolution details issues Feb 2, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

output image resolution details issues #126

output image resolution details issues #126

bjohn22 commented Feb 2, 2025 •

edited

Loading

bjohn22 commented Feb 2, 2025

output image resolution details issues #126

output image resolution details issues #126

Comments

bjohn22 commented Feb 2, 2025 • edited Loading

bjohn22 commented Feb 2, 2025

bjohn22 commented Feb 2, 2025 •

edited

Loading